Get in touch

9io.ai / Blog

What an AI product development company does differently in 2026

Capable LLMs are plentiful in 2026. What an AI product development company adds is the engineering around them, shown with numbers from our own builds.

Key takeaways

  • Capable models are plentiful in 2026, and the engineering around them is what separates a demo from a product.
  • Choose a model and a reasoning effort for each step, and keep a second provider wired in.
  • Check every model answer in code, and keep a rules-based path for when the model or judge fails.
  • Our in-app assistant now shows first text at about 3.2 seconds, down from 20 to 30.
  • Log spend per model and per job, and enforce each job’s daily budget in code.

Capable language models are plentiful in 2026. Our comparison of 23 models from 11 labs found input prices from $0.10 to $10 per million tokens as of 8 October 2026, and no model leading every benchmark. A convincing demo on top of any of them is quick to build. What an AI product development company adds is the engineering that makes the demo hold up with real users, and the evidence for that engineering is numbers from products in production.

We call 9io an AI-first product studio, and a label like that is easy to claim. This post shows what it means, using field notes from products we build and run. On an AI study platform, an in-app assistant used to take 20 to 30 seconds to start answering, and it now shows its first text at about 3.2 seconds. On a fashion commerce platform, an AI stylist’s median response time fell from 30.1 s to 15.5 s on live traffic. Both gains came from changes to how the model was called and what surrounded it.

The sections below cover model choice, evals, latency, cost, MCP and when not to use a model, then a checklist for judging any AI product studio and how an engagement with us starts.

What changes between an LLM demo and an LLM product

As of 8 October 2026, GPT-6 Luna and Claude Haiku 5.5 both cost $0.10 per million input tokens and $0.50 per million output tokens. Within one model, moving from low to maximum reasoning effort multiplied the cost per task by about 11 on Claude Opus 5.5 and about 16 on Claude Sonnet 5.5. Six labs changed prices or replaced models between July and October, so a model choice needs rechecking every month.

In LLM app development, the parts that last are the ones built around the model, and a demo needs none of them. The table lists what changes, with the field note where each concern shows up in our builds.

Concern In a demo In a product Field note
Model choice One model, default settings, every step A model and reasoning effort per step Content pipeline: reasoning was 187,152 of 572,209 output tokens in one build
Providers One provider Several, with a fallback route when latency degrades AI stylist: calls moved off Azure OpenAI to avoid a slow tail
Answer checks A person reads the output Two reviews, a judge, checks in code and a rules-based fallback Research pipeline: confidence capped at 60 when the reviews disagree
Regressions A few prompts tried by hand A test suite run on every change AI stylist: 18 of 47 cases passing before, 47 of 47 after
Missing output Errors appear on screen Planned output compared with shipped output Content pipeline: 26 of 51 diagrams shipped, report said 0 dropped
Latency Fine on the demo laptop First byte, first text, p50 and p90 tracked on live traffic Assistant: 20 to 30 s down to about 3.2 s to first text
Spend Nobody checks the bill A ledger row per call and a daily budget per job in code Vision bill: about $1,100 on 419,057 unneeded calls
Distribution A chat box in the app The same tools in Claude, ChatGPT, Cursor and VS Code MCP server: five failures, almost all in OAuth and protocol versions
Tool permissions The model can call any tool Limits for outside callers enforced on the server MCP server: order-placing actions off for third-party clients
Whether to use a model A model for everything Rules and code wherever they give the right answer Research pipeline: a rules-only vote decides when no model key is set

How we choose models when we build AI products

We choose a model and a reasoning effort for each step, because steps differ in how much judgement they need. A step that only reshapes text that already exists, such as filling a fixed schema from a finished explanation, is mechanical. Reasoning tokens never appear in the response, but they are billed as output tokens.1 In one build of the course-building pipeline on the AI study platform, reasoning took 187,152 of 572,209 output tokens, about 3.8 for every token of prose that shipped. We moved its mechanical steps to low reasoning effort.

Our rule is to run mechanical steps on a small model at low effort and keep larger models for the steps that need judgement. The AI stylist’s compose step, which writes from material it has already been handed, now runs with reasoning off, and the study platform’s assistant runs at low effort. Traffic from the fashion commerce platform’s embeddable widget goes to a cheaper model. 9io Alpha’s research pipeline, where the reviews and the judge all make judgements, runs every call at medium effort with output capped at 8,000 tokens. Lower effort trades depth of reasoning for speed, so we compare output before and after each change.

The products we build also mix providers. Across the work in our field notes we use OpenAI’s models, Google’s Gemini, open models such as FashionSigLIP, GroundingDINO and SAM 2, and Anthropic’s Claude through MCP clients. The AI stylist alone calls OpenAI’s API and retrieves from Azure AI Search. A second provider wired in gives a call somewhere to go when its route degrades, and the stylist’s model calls moved from Azure OpenAI to OpenAI’s API after we saw a slow tail on Azure.

Model families matter for reviews too. A study of more than 350 models found that when two models both answered wrongly on one leaderboard, they gave the same wrong answer 60% of the time, more often when they shared a provider or architecture.2 9io Alpha’s two reviews still run on one model, with a path for a second provider switched off, and moving one review to another family is on our list of changes.

Evals, an LLM judge and checks in code

Our AI features are tested on every change, and their pipelines can be traced. That is what makes the model and effort choices above safe to revisit. On the AI stylist, a quality pass ran alongside the latency work. The regression suite went from 18 of 47 cases passing to 47 of 47, and answers judged bad in review fell from 21 of 116 to 1 of 116.

In 9io Alpha’s research pipeline, a model’s answer is checked by other model calls and then by code. Two differently prompted reviews of the same evidence argue for and against, and a judge reads both, plus a vote from plain rules, and makes the call. Code then checks the verdict against an allowed list, clamps numbers to their ranges, requires a written rationale and caps confidence at 60 when the reviews disagree. If the judge times out or returns something unusable, a fixed merge of the votes decides, and each response records which path produced it.

In a 2023 study, strong LLM judges agreed with human preferences more than 80% of the time, about as often as humans agree with each other, and showed position, verbosity and self-enhancement biases.3 Our judge always sees the two reviews in the same order, and we haven’t yet tested whether swapping them changes its verdicts.

A judge can also fail without anyone noticing. In one run of the course-building pipeline, the judge left 27 of 35 sections unreviewed, and the counters showed 774 deadline exhaustions and no throttle retries. We raised its time budget from 90 s to 420 s and split judge errors into unreachable and unparseable, since each needs a different fix. A section in either state counts as unreviewed.

The same pipeline taught us to compare planned output with shipped output. One build planned 51 diagrams and shipped 26, planned 44 simulations and shipped 13, and its report said 0 dropped. A concurrency limiter created per call had let every module hit one deployment at once, and work that timed out after 429 responses was never counted. After five fixes, a later build had 33 of 33 sections at full length, against 18 of 38 before.

Cutting the wait before the first word

By Jakob Nielsen’s measure, about 1 second keeps a user’s flow of thought uninterrupted, and about 10 seconds is the limit for keeping their attention.4 Our in-app assistant was past the second limit, taking 20 to 30 seconds to its first word.

Each turn chained three slow steps before the model could answer. A forced tool call fetched context, the tool ran in the browser and cost a round trip, and then the course content was read. Now the current page, the highlighted passage and the chosen skill go straight into the instructions, along with the course text, which was already loaded for the access check. Tools are forced only on turns that must end in an action, and reasoning effort is low. First text appears at about 3.2 seconds and the first byte at about 0.8 seconds.

The AI stylist runs a chain of model calls on every turn, so their times add up. Its median response time was 30.1 s across 394 real turns before tuning, and 15.5 s over four days of live traffic after it. The p90 fell from 40.3 s to 20.6 s. Four changes went in at once. Reasoning was turned off for the compose step, slow calls were hedged, the calls moved to OpenAI’s API, and retrieval now stops once it has enough results. The figures cover all four, so they don’t show what each contributed.

A hedged request sends a copy of a slow call after a delay and keeps whichever answer arrives first. It helped while the tail was long. Once the other changes had shrunk the tail, a later sample showed hedging fire 142 times, and the duplicate finished first in only 4 of them. A hedge can beat a slow server, but it can’t shorten an answer that is slow because it is long, since the copy has the same tokens to generate and starts later.

Keeping LLM spend visible and capped

An LLM bill usually comes to more than list price times visible tokens, because of reasoning tokens, long-prompt surcharges, cache-write fees and differences between tokenizers. So we record spend where it happens. In the course-building pipeline, every model call now writes a ledger row with the model, the kind of work, the token counts by type and a status. Grouped by model and by kind, the rows show which step to change, and the status column shows calls that produced nothing.

On the fashion commerce platform, a background job re-classified items it had already classified. A variable named for a two-day window held 730 days, and nothing skipped finished items, so the job made 419,057 vision calls it didn’t need, for about $1,100. At about a quarter of a cent a call, no single call stood out. Provider rate limits cap throughput, so a job running steadily below them raises no error, and provider spend controls act late.

These points from the field note’s checklist would have caught this job early.

  • The selection skips items that already have a result from the current model and prompt version.
  • The job estimates its cost before the first paid call, and a run over a threshold needs an explicit flag from a person.
  • Each job has a daily budget enforced in code, with an alert at half of it.
  • Batch jobs run in their own project or workspace, with a lower provider-side limit than user-facing features.

The embeddable widget is now capped at 6,000 generations a day, so its worst day has a known cost.

Meeting users in Claude, ChatGPT, Cursor and VS Code

An MCP server lets people use a product’s tools from the AI clients they already work in. We’ve built remote MCP servers for two products, and one server answers Claude, ChatGPT, Cursor, VS Code and command-line clients from a single URL.

Five things broke on the way, almost all in OAuth and protocol-version handling. Cursor’s registration was refused over one redirect URI on its own scheme, command-line clients couldn’t find the OAuth discovery documents, requests on the newer protocol version got 400s, and sign-in worked with only one of the product’s login methods. ChatGPT’s custom GPTs accept at most 30 operations per OpenAPI schema, so the REST mirror for them ships as two files. The fashion commerce platform’s server now passes a 42-check end-to-end test.

When a third-party client calls the fashion commerce platform’s shopping assistant through its MCP server, the assistant runs with its order-placing actions switched off, and that rule is enforced on the server. Tool annotations such as a read-only hint only describe a tool to the client, and the MCP specification says clients must treat them as untrusted unless they come from trusted servers.5

When a form or a database query beats a model

Some features don’t need a model, and we say so when that’s the case. If a form or a database query solves the problem, use that. It gives the same answer every time and is tested like any other code.

The same thinking applies inside AI features. In 9io Alpha’s research pipeline, a vote from plain rules sits beside the two model reviews. When no model key is configured, no model call is made and the rules vote becomes the verdict. Numbers the judge proposes are clamped by rules in code before anything acts on them.

For a bounded judgement made many times, a smaller tool may fit better, and our comparison of decision models, LLMs and classifiers gives a rule. Fine-tune a classifier when the categories are stable and labelled examples number in the thousands, and use a decision model when categories change often and answers must come within about a second. Use an LLM with structured output when the answer needs words, arithmetic or several steps of reasoning. Policy limits, and anything code can compute exactly, stay in code.

How the numbers in this post were measured

Figures about our own work come from the linked field notes, and market figures from the model comparison, read on 8 October 2026. We took no new measurements for this post.

  • Assistant latency comes from production. The field note gives no percentiles, sample sizes or split by change.
  • Stylist latency compares 394 real turns before tuning with four days of live traffic in late September 2026. The hedging counts come from a later live sample, and the quality figures from a 47-case suite and a review of 116 answers.
  • Pipeline figures come from single builds, with 36 sections for planned against shipped, 38 and 33 sections for before and after, and a separate 35-section run for the judge. The token counts are from one build.
  • Spend is an exact count of 419,057 calls attributed to one job, with the cost rounded to about $1,100.
  • 9io Alpha logs which path produced each verdict but doesn’t aggregate it yet, so we can’t say how often the fallback runs.

Questions to ask an AI product development company

We’d put these questions to any studio before hiring it to build an AI product, us included.

  • Production numbers. Ask for figures from a product with real users, with the sample size beside each. A replay of 12 requests is good for comparing versions but can’t pin down a p90.
  • A model per step. Ask which model and reasoning effort each step uses, and what the test suite showed when that changed.
  • The failure path. Ask what happens when the model or the judge fails, and whether the product records which path produced each answer.
  • Missing output. Ask how they’d notice work that was planned and never shipped.
  • Latency. Ask for p50 and p90 from live traffic, with first byte and first text tracked separately.
  • Spend. Ask to see the ledger per job and model, the daily caps, and what stops a runaway job before the provider’s limit does.
  • Outside callers. If the product exposes an MCP server, ask what a third-party client can do through it and where that’s enforced.
  • No model. Ask which parts they’d build with rules, queries or forms.
  • People and code. Ask who would do the work, and whose repository the code lives in from the first commit.

Commercial questions, such as hours per person and who pays the model bills, are in our cost model for AI apps.

How an engagement with 9io starts and runs

Starting takes one email describing the product and what you want the AI to do. The reply comes from the person who’d do the work. If the build goes ahead, the proposal names everyone who’d work on it, with the hours and the rate for each person. Work usually starts within a couple of days of agreeing it, and if we can’t start that quickly, we say so.

Once work starts, these terms apply.

  • We bill only for the hours we work, with no management fee and no hidden extras.
  • Invoices go out twice a month for work already done, and 10% of each waits until you accept the milestone. Each checkpoint gets seven working days for review.
  • You get a working environment and a login near the start of the build, and a short written update every Friday.
  • Code, history and infrastructure definitions live in your repository from the first commit, and you own the work as you pay for it.
  • Defective delivered work is fixed under a 90-day warranty.
  • You can scale up, scale down or stop at a month boundary with 30 days’ notice, with no retainer to unwind.
  • Uptime and error alerting before launch, tested backups, and documentation with a runbook come as standard.

More on how we approach AI builds is on our AI engineering page.


  1. OpenAI API documentation, “Reasoning models”, https://developers.openai.com/api/docs/guides/reasoning ↩

  2. Elliot Kim, Avi Garg, Kenny Peng and Nikhil Garg, “Correlated Errors in Large Language Models”, ICML 2025, https://arxiv.org/abs/2506.07962 ↩

  3. Lianmin Zheng et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”, NeurIPS 2023 Datasets and Benchmarks Track, https://arxiv.org/abs/2306.05685 ↩

  4. Jakob Nielsen, “Response Times: The 3 Important Limits”, Nielsen Norman Group, adapted from Usability Engineering (1993), https://www.nngroup.com/articles/response-times-3-important-limits/ ↩

  5. Model Context Protocol specification 2026-07-28, “Tools”, https://modelcontextprotocol.io/specification/2026-07-28/server/tools ↩

Frequently asked questions

What does an AI product development company do?

It builds the engineering around the models. That means choosing a model and reasoning effort for each step, testing answers on every change, keeping latency and spend within budget, and saying where a model isn’t needed at all.

What is an AI-first company?

A company that builds its products around AI from the start. For an AI product studio, the useful test is whether it can show production numbers for quality, latency and cost, and whether it tells you when a feature doesn’t need a model.

How do you choose which LLM to use for each step of an app?

Shortlist models by price and context needs, then test them on your own inputs. Run mechanical steps, such as filling a template, on a small model at low reasoning effort, keep larger models for steps that need judgement, and compare output before and after each change.

How can I make an LLM app respond faster?

List everything a turn waits on before the first token, and remove the steps that fetch what you already have. Our assistant went from 20 to 30 seconds to about 3.2 seconds to first text after its context moved into the prompt, tools were forced only for actions and reasoning effort was set to low.

How do you control LLM API costs in production?

Write a ledger row for every paid call with the job, model, units and cost, and enforce a daily budget per job in code. Provider spend limits act late, so treat them as an outer layer and give batch jobs lower limits than user-facing features.

When should a product not use AI?

When a form, a database query or a fixed rule gives the right answer. Those give the same answer every time and are tested like any other code, so keep models for the judgements code can’t make.

Work with us

Building something like this?

9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.