Artificial Intelligence Models: How to Interpret and Evaluate Them in a Business Setting

On September 28, 2026, Anthropic unveiled Claude Sonnet 5.5.

In the announcement for the Claude Sonnet 5.5 (*) reposted by Tom’s Hardware, one number stands out: on Terminal-Bench 4.0, a programming test, the new model scores 70.6%, compared to 10.3% for its predecessor.

(*) Source: www.tomshw.it/hardware/con-claude-sonnet-55-anthropic-alza-ancora-lasticella-dellia-ecco-tutte-le-novita

In the comments section below the article, a reader wonders whether the old model was really that bad, or if something doesn’t add up.

That’s the right question. Every month, new artificial intelligence models are released, each accompanied by score tables, acronyms, and percentages.

For those who work in a company, the challenge is rarely finding information; it is figuring out which pieces of information matter for their decisions.

This article aims to provide some clarity, with a short glossary for those who use AI on a daily basis and a set of criteria for those who need to decide whether to adopt it.

The Same News, Two Perspectives

A lot of shared information circulates within the company: a report, an announcement, a demo seen at a trade show.

Everyone interprets them through the lens of their own role, and everyone seeks different answers. This is a source of strength: the best decisions emerge when those different perspectives come together around the same table.

Who Uses AI Every Day

An employee who writes reports, analyzes data, or prepares presentations reads the ad with a practical question in mind:

  • What can I do better—or faster—with this tool?

He needs the right words to understand the technical data sheets and to explain to his colleagues what works and what doesn’t.

Who decides whether to adopt it?

The entrepreneur reads the same ad, but with different questions.

  • Is it worth it?
  • How reliable is it when it comes to our processes?
  • What will change for people?

Behind that last question often lies a real fear: the fear of having to lay off staff.

A well-planned adoption process starts right here: first deciding how to use the time freed up from repetitive tasks, which can be redirected toward higher-value tasks or returned to people in the form of more sustainable schedules.

What Are Artificial Intelligence Models?

An artificial intelligence model is the engine behind tools such as ChatGPT, Claude, and Gemini: a system trained on vast amounts of text, images, or data that has learned to recognize patterns and generate responses.

The application you use every day is the body; the model is the engine, and the same engine can be used in different products.

If you’d like to learn more about the difference between artificial intelligence and machine learning, you’ll find it explained in a dedicated article.

General Models and Specialized Models

General models can write, summarize, translate, and analyze documents and images.

Specialized models are trained or adapted for a specific task: detecting defects on a production line, reading contracts, and forecasting demand.

Manufacturers often offer multiple versions of the same product line: a more powerful one for complex tasks and a faster, more affordable one for everyday use.

This is the case with Claude Opus and Claude Sonnet, which are designed to complement each other.

Why are new versions released every few months?

Competition among manufacturers is fierce, and each new version promises greater capacity, faster speeds, or lower costs. For a company, this has only one practical implication: the choice of a model is subject to change.

It’s a good idea to establish processes that allow you to change it without having to start over from scratch.

Terms for Using AI

These are the terms that users encounter in their daily work. Knowing them helps you ask the tool for the right thing and talk about it with colleagues using the same vocabulary.

Generative AI, language models, and prompts

Generative AI produces new content: text, images, audio, and code.

Large language models (LLMs) are the most common type: they have learned from human languages to predict, word by word, the most plausible continuation of a text.

The prompt is the request we send to the model: the clearer it is about the context and the expected result, the more useful the response will be.

NLP and Computer Vision

Natural language processing (NLP) is the branch of AI that understands and generates language: it enables an assistant to respond to customers or sort incoming emails.

Computer vision does the same with images: it recognizes objects, reads labels, and identifies defects. In a fashion company, it can check the quality of a fabric; in a warehouse, it can read package codes.

RPA and AI agents

RPA (robotic process automation) automates repetitive tasks by following set rules, such as copying data from a management system to a spreadsheet.

An AI agent goes one step further: it is given a goal and decides on its own the sequence of actions needed to achieve it, using multiple tools.

The difference is similar to that between following a recipe to the letter and a chef who adapts the dish to the ingredients on hand.

Criteria for evaluating it

It’s the language used in advertisements and technical specifications that makes it difficult to read news stories like the one we started with.

Benchmarks, Scores, and Test Versions

A benchmark is a standardized test: a set of tasks that are the same for all models, allowing them to be compared. The score indicates the percentage of tasks completed correctly.

Tests also have versions, and a comparison only makes sense when comparing the same version.

The announcement for Sonnet 5.5 includes three programming benchmarks:

  • Terminal-Bench 4.0: 70.6% score
  • FrontierCode 1.1: 46.2% result
  • CursorBench 4.0 result: 55.5%.

Token, context window, and cost per task

A token is the unit that the model uses for reading and writing: a word fragment.

Services are paid for using tokens, distinguishing between input tokens—that is, what we send—and output tokens—that is, what the model produces.

The context window is the amount of text that the model considers at one time, much like a person’s working memory.

Per-task pricing shifts the focus from the unit price to the cost of a complete task: a more expensive token-based model can actually cost less if fewer tokens are used to achieve the result.

That’s what Anthropic claims for the Sonnet 5.5: up to 30% lower power consumption compared to its predecessor.

Level of Reasoning and Long-Form Tasks

Some models let you choose how long to think before answering: this is the level of reasoning, also known as “effort.”

More reasoning generally leads to better answers to complex problems, with more time and more tokens.

That is why the 46.2% Sonnet 5.5 score on FrontierCode—achieved at the highest setting—should be interpreted with the understanding that the configuration may differ in everyday use.

Long-term tasks, also known as agentic tasks, measure the ability to complete multi-step activities.

The test on *Pokémon Red*, completed by Sonnet 5.5 using only screenshots from the game, serves precisely this purpose: to verify whether the model interprets what it sees and stays on course over a long period of time.

How do you interpret an AI benchmark score?

An AI benchmark score indicates the percentage of standard tasks that a model solves correctly.

It should be read with four checks in mind:

  • who conducted the test,
  • which version was used,
  • at what level of reasoning
  • how much those tasks resemble the company’s actual work.

Without these checks, the number simply compares the models and says little about their usefulness.

Who took the test?

The results published in advertisements are often measured by the manufacturer itself.

This doesn’t mean they’re inaccurate; it suggests waiting for independent verification—such as public rankings and third-party tests—before treating them as fact.

Had the model already seen the questions?

The models learn from enormous amounts of publicly available text.

If benchmark questions are circulating online, they can end up in the training data: it’s like a student who has seen the test in advance.

That is why the tests are updated to new versions, and a very large jump—such as the one from 10.3% to 70.6%—should be interpreted in conjunction with the results from other tests.

How much does the test resemble your job?

A programming benchmark tells software developers a lot, but it tells contract administrators and hotel staff very little.

The most useful test is the one that resembles your own process.

In the announcement, Anthropic cites an internal test in which the model produced a ten-slide presentation deemed usable without modifications: an indication of its suitability for office work, though based on tasks selected by the developer.

What criteria should you use when choosing an AI model for your business?

When choosing an artificial intelligence model, it’s best to start with the process you want to improve, rather than with the rankings.

There are six main criteria: performance in a real-world test using the organization’s own data, cost per completed task, data confidentiality, integration with existing systems, known limitations of the model, and the possibility of replacing it in the future.

The trial run for your process, before signing the contract

A pilot test lasting just a few weeks on a limited-scope process speaks louder than any ranking: you select ten or twenty real-world cases, compare two or three models, and evaluate the results together with the people who do that work every day.

In the FARO by Factory method, this phase is called “Focus”: defining the objective, process, and metric before selecting the tool.

It’s the starting point for anyone who wants to implement artificial intelligence in their company and achieve measurable results.

Data, Confidentiality, and Investment

Before comparing performance, it’s important to know where the data is processed, whether it’s used to train the model, and under what contractual terms.

The investment depends on specific factors: usage volume, process complexity, integrations with ERP and CRM systems, and employee training. The model accounts for a portion of the total cost—often not even the largest portion.

Limitations and Biases of Models

Every model has its limitations: it can generate information that is plausible but incorrect, reflect distortions in the training data—known as biases—and behave differently in Italian than it does in English.

That’s why you always need someone to review and approve the result.

Adopting AI in an ethical manner also means this: determining who is in control and informing customers and employees where it is used.

Five Questions to Ask Before Trusting a Benchmark

A quick checklist to keep on hand for your next ad campaign—useful for both AI users and decision-makers.

The Five-Question Checklist

  1. Who measured the result: the manufacturer or a third party?
  2. Which version of the test, and at what level of reasoning?
  3. Is the comparison made with models in the same category, using the same version of the test?
  4. Do the tasks in the test resemble the processes at my company?
  5. How does the model perform when tested with my data, and how much does each completed task cost?

From Reading to Decision-Making

These five questions help you analyze a job posting.

The text in this article is meant to help you read the ads. Decisions are made elsewhere.

In projects where I work with companies looking to implement artificial intelligence, models, tokens, and benchmarks are the last topics to come up in the conversation.

The first questions concern the value that AI can create for entrepreneurs, companies, and their employees:

  • Which processes are repetitive?
  • where workflows can become simpler
  • what information gets lost along the way
  • which should be shared and which should rightfully remain confidential.

Anyone can use AI to create a great presentation, and there are excellent tutorials available online.

Introducing it into the company in an ethical manner, with a medium- to long-term vision, is a systematic process: it starts with people and processes, and the tool is chosen only at the end.

Want to learn how to implement AI in your company? Request a free consultation.

Frequently Asked Questions About Artificial Intelligence Models

What are the main models of artificial intelligence?

Among the most widely used general-purpose models are OpenAI’s GPT family, Anthropic’s Claude, and Google’s Gemini, alongside open-source models such as those from Mistral and Meta.
The choice depends on the task, the language, data processing, and integrations with tools already in use at the company.

What is an AI benchmark?

An AI benchmark is a standardized test: a set of tasks that are the same for all models, allowing their performance to be compared.
The result is usually expressed as a percentage of tasks completed correctly and is meaningful only when comparing models tested on the same version of the benchmark.

Are AI benchmarks reliable?

They are reliable as a benchmarking tool, but less so as a measure of business value.
It’s a good idea to check who conducted the test, whether the model might have already been familiar with the questions, and how closely the tasks resemble your own processes. Testing on your own data remains the most reliable way to verify results.

What does “token” mean in artificial intelligence?

A token is the unit used by a model to read and write text: a fragment of a word.
AI services are billed based on input and output tokens, and the context window indicates how many tokens the model can process at a time.

Which paid artificial intelligence solution is the best choice for a business?

It depends on the process you want to improve. It’s a good idea to compare two or three solutions through a pilot test using real-world cases, evaluating the quality of the results, the cost per completed task, data management, integration with existing systems, and support.
The overall ranking is a starting point, not an answer.

What metrics should be used to evaluate AI in a company?

The most useful metrics are those related to the process: time spent per task, errors and rework, volume handled, and customer and employee satisfaction.
These metrics should be measured before and after the introduction of AI, using the same process and over a period long enough to observe a stable trend.

Will artificial intelligence replace employees?

AI replaces tasks, not people. Companies that adopt it systematically decide in advance how to use the time freed up: higher-value tasks, training, customer relations, and, in some cases, more sustainable work schedules.
Involving employees from the very beginning is essential for the change to succeed.

What factors determine the cost of an artificial intelligence project?

It depends on the volume of use, the complexity of the process, integrations with the ERP and CRM systems, the quality of the source data, and staff training.
The cost of the model is only part of it: the total investment is determined based on the project, after the objectives and processes have been clarified.

About this item
Share this article
If you need support, or want to understand how we can help your Company contact us now:
Would you like to receive information?
Fill out the form
This site is protected by reCAPTCHA and Google: Privacy Policy e Terms of Use.