Get in touch

9io.ai / Blog

How much does it cost to build an AI agent or app in 2026?

How much does it cost to build an AI agent? A cost model that shows its working, from build hours and market rates to monthly running costs per 1,000 users.

Key takeaways

  • Every development quote is hours times a rate, so ask to see the hours.
  • Our cost model puts a first production AI agent at 330 to 980 hours, or $32,000 to $245,000.
  • On a mid-tier model, tokens cost far more than hosting, and a small model can be 20 times cheaper.
  • Evals, security review, spend limits and EU marking duties are easy to leave out of a quote.
  • METR’s 2026 data shows small, uncertain speedups from AI coding tools, so don’t assume a big discount.

On the cost model in this post, the answer to “how much does it cost to build an AI agent” is about 330 to 980 engineering hours for a first production release. At 2026 market rates of $95 to $250 an hour, that comes to roughly $32,000 to $245,000. Running it for 1,000 users at moderate use then costs about $900 to $14,300 a month, upkeep included. The number moves most with how many systems the agent touches, how much testing it needs before launch, which model tier it runs on and whose hourly rate you pay.

Most published AI app development cost figures for 2026 are ranges without the working. GoodFirms’ 2026 guide, based on input from 267 mobile app development companies, puts AI app development at $50,000 to $300,000 and up.1 The MVP price ladders that AxonBuild collected from six development companies in September 2026 have middle tiers that together span $15,000 to $100,000.2 Neither tells you where your product falls.

This post shows the working. Rates come from public sources, effort comes from labelled assumptions you can replace, and running costs are computed from providers’ price lists as of 8 October 2026. We build AI products for clients and work by the hour, so this is also the arithmetic we’d want a buyer to run on any quote, ours included.

Every quote is hours times a rate

Custom software is priced in hours even when the contract has a fixed price. A fixed price is an estimate of hours multiplied by a rate, plus a margin for risk. AxonBuild’s roundup makes the same point about MVP tiers, where a name such as “Standard MVP” stands in for a number of hours.2 So the cost model has one formula.

build cost = hours × blended hourly rate × (1 + contingency)

There is plenty of public data on rates and very little on the hours an AI build takes. GoodFirms gives timelines, such as 9 to 12 months for an advanced app with AI or machine-learning features, but a timeline doesn’t say how many people work on it.1 We found no survey that reports engineering hours for a RAG assistant, a voice agent or an agentic workflow. The hours below are therefore assumptions, built up from work packages, and you should replace them with the hours in any quote you get.

The tables carry no contingency. GoodFirms advises adding 25% to 35% to an app budget for hidden costs, although that allowance also covers app-store fees and marketing.1

What an hour of AI app development costs in 2026

The US Bureau of Labor Statistics (BLS) puts the median pay of US software developers at $135,980 a year in May 2025, which is $65.38 an hour over a 2,080-hour year.3 In June 2026, wages and salaries were only 69.2% of employer costs for private-industry professional and related workers. The rest went on paid leave, supplemental pay such as bonuses, insurance, retirement plans and legally required benefits.4 Dividing $73.69 of total cost per hour worked by $50.98 of wages gives a multiplier of 1.45. A median developer on payroll therefore costs about $94.50 per hour worked, and one at the 90th percentile of pay ($214,670) about $149.

Rates as of 8 October 2026:

Who Rate per hour Source
Median US software developer on payroll, with benefits and payroll taxes about $94.50 BLS34
90th-percentile US software developer on payroll, same basis about $149 BLS
Freelancers $25 to $150 GoodFirms1
Development agencies $75 to $250 GoodFirms
Senior developers in the US, by platform $100 to $200 and up GoodFirms
Fractional senior software engineers median $175, middle half $130 to $200 Go Fractional5
Fractional CTOs median $200, middle half $150 to $250 Go Fractional6

Some notes on the table:

  • The payroll figures leave out recruiting, management, equipment and the weeks it takes to hire, and they assume the person stays after the build.
  • GoodFirms’ rates come from mobile app companies. Its senior rates for India ($25 to $70, depending on platform) and Poland ($45 to $100) are far lower, and the arithmetic works the same way.
  • Go Fractional’s figures cover job posts and candidate profiles from the 90 days to 8 October 2026. The senior engineer benchmark rests on 11 posts and 133 profiles, so treat it as indicative. The CTO benchmark uses 18 posts and 930 profiles, and puts a typical engagement at 15 hours a week, or $9,000 to $15,000 a month.

We price each build at $95 an hour (a median developer on payroll), $175 (the median fractional senior engineer) and $250 (the top of GoodFirms’ agency range).

How much does it cost to build an AI agent, in hours and dollars

The cost model prices three reference builds. Each is a first production release for one company, in one language and one region, with a web interface.

  • RAG assistant over company documents. Staff ask questions in a web chat and get cited answers from about 20,000 documents in two sources, such as a shared drive and a wiki. People sign in with the company’s single sign-on and only get answers from documents they can open. The eval set has 300 questions with known good answers.
  • Voice agent. It answers a phone line, handles common questions from a small knowledge base, and looks up and changes bookings through two or three APIs. It says it is an AI at the start of every call and hands over to a person with a summary when it should. The eval set is 100 simulated calls, plus latency checks.
  • Agentic workflow. It picks up requests from a queue, reads records in four internal systems through their APIs and makes changes. Anything above a risk threshold waits for a person to approve it. Every action is logged, writes are safe to repeat, and each task has step and spend limits. The eval set is 200 tasks with expected end states, including prompt-injection attempts.

Hours per work package. Every figure in this table is an assumption:

Work package RAG assistant Voice agent Agentic workflow
Discovery, architecture and data audit 24–40 30–50 40–60
Integrations (document sources; telephony and booking APIs; internal systems) 60–120 100–200 120–240
Core AI logic (retrieval and answers; real-time voice; agent loop) 70–140 80–160 60–120
Access control, approvals and hand-off 40–80 30–60 60–120
User interface 40–80 40–80 40–80
Eval set and harness 40–80 60–120 80–160
Guardrails, observability and spend limits 20–40 30–60 60–120
Security review, deployment and runbooks 40–80 40–80 40–80
Total hours 334–660 410–810 500–980
Weeks with two engineers at 30 hours a week each 6–11 7–14 8–16

Integrations get the widest ranges because their cost depends on the other system’s API, which nobody has seen in detail before discovery. Eval work grows with the ways a build can fail. A RAG assistant can give a wrong answer, a voice agent can also mishear or talk over the caller, and an agent can also take a wrong action. That last risk is why the agent’s approvals, guardrails and evals carry the most hours.

Build cost at the three rates:

Build Hours At $95 At $175 At $250
RAG assistant 334–660 $31,730–$62,700 $58,450–$115,500 $83,500–$165,000
Voice agent 410–810 $38,950–$76,950 $71,750–$141,750 $102,500–$202,500
Agentic workflow 500–980 $47,500–$93,100 $87,500–$171,500 $125,000–$245,000

The cost to develop an AI agent that acts in other systems is the highest of the three, because every write needs permissions, an approval path and tests for the actions it must never take. For an AI MVP, development cost sits toward the low end of each range, since an MVP usually trims the integrations, the interface and the eval set.

The rate moves the total as much as the scope does. On the same hours, $250 costs 2.6 times as much as $95. At $175 to $250 an hour, the reference builds overlap the band GoodFirms gives for an advanced app with AI features, $100,000 to $250,000 and up.1

Monthly LLM spend per 1,000 users

On a mid-tier model, most of the running cost is model calls. This cost model reuses two request shapes from our LLM API price comparison, which prices them on 13 models, so you can swap in another model’s figure.

  • A RAG question is a chat answer, with 8,000 input tokens of instructions and retrieved context and a 500-token reply. Claude Sonnet 5.5 costs $2 per million input tokens and $10 per million output tokens, so a question costs 8,000 × $2 / 1M + 500 × $10 / 1M = $0.021.7 On Claude Haiku 5.5, at $0.10 and $0.50, it costs $0.00105.
  • An agent task is five agent steps. Each step sends 50,000 input tokens, 45,000 of them from a warm prompt cache, and gets 2,000 back. With cached input at $0.10 per million, a Sonnet 5.5 step costs 45,000 × $0.10 / 1M + 5,000 × $2 / 1M + 2,000 × $10 / 1M = $0.0345. A task costs $0.1725, or $0.00975 on Haiku 5.5.
  • A voice minute runs on OpenAI’s GPT-Live-1, which costs $0.05 per minute of session, billed per second. Calls to the backend model it delegates to are charged at that model’s normal rates.89 We assume one backend call a minute to GPT-6.1 Sol, sized like the chat answer, which adds $0.021 and makes $0.071 a minute.

GPT-6.1 Sol and GPT-6 Luna cost the same as Sonnet 5.5 and Haiku 5.5 for these shapes.8 The usage levels are assumptions. A light, medium or heavy user asks 20, 100 or 400 questions a month, talks for 10, 30 or 120 minutes, or runs 10, 50 or 200 tasks.

Model spend per month for 1,000 users, as of 8 October 2026:

Build and model Light Medium Heavy
RAG assistant on Sonnet 5.5 or GPT-6.1 Sol $420 $2,100 $8,400
RAG assistant on Haiku 5.5 or GPT-6 Luna $21 $105 $420
Voice agent on GPT-Live-1 with GPT-6.1 Sol $710 $2,130 $8,520
Agentic workflow on Sonnet 5.5 or GPT-6.1 Sol $1,725 $8,625 $34,500
Agentic workflow on Haiku 5.5 or GPT-6 Luna $97.50 $487.50 $1,950

The model tier changes the bill by 18 to 20 times. Whether the small model is good enough is a quality question. Artificial Analysis’s Intelligence Index scores Sonnet 5.5 at 56 and Haiku 5.5 at 43,10 and only your eval set can say which one your task needs. That is a good reason to build the eval set before you choose a model.

Reasoning tokens are billed as output and aren’t in these shapes. With 500 reasoning tokens per answer, a Sonnet 5.5 question costs $0.026 instead of $0.021. Prices also move during a project. Gemini 3.8 Flash costs $0.75 per million input tokens and $3.75 per million output tokens until 31 December 2026, and twice that from 1 January 2027.11 The same chat answer goes from $7.88 to $15.75 per thousand.

For voice, GPT-Live-1 charges for session time, so an open session costs the same whether or not anyone is speaking. OpenAI’s token-priced Realtime models send the whole conversation to the model for each response, so later turns cost more than early ones.12 If the agent answers a phone line, Twilio adds $0.0085 a minute for inbound calls to a local US number. Media Streams, which carries the call audio to your server, adds $0.0044 a minute.13

Hosting, vector storage, evals and observability

On a mid-tier model, the rest of the monthly bill is smaller than the model calls. Each line is an assumption priced from the provider’s list as of 8 October 2026.

Monthly line RAG assistant Voice agent Agentic workflow What it assumes
Hosting on Render $161 $161 $161 Pro workspace, two web instances and one worker with 1 CPU each, Postgres with 4 GB of RAM and 20 GB of storage
Vector storage on Pinecone $50 $20 $0 Standard plan minimum; Builder plan flat fee; no vector store
Observability on Langfuse $61 $31 $69 Core plan; 5 units per question, 4 per voice minute, 12 per task
Eval runs $126 $447 $387 10 full runs a month, each case scored by an LLM judge
Phone line on Twilio – $388 – 30,000 inbound minutes with Media Streams, one local number
Total $398 $1,047 $617

How each line was priced:

  • Hosting. Render charges $25 a month for the Pro workspace, $25 for each 1-CPU, 2 GB instance, $55 for Postgres with 4 GB of RAM and $0.30 per GB of database storage.14
  • Vectors. Assume the 20,000 documents average 7,500 tokens, which makes about 300,000 chunks of 500 tokens. With 1,536 dimensions and 1 KB of metadata each, they take about 2.1 GB. Pinecone’s rates put that at $0.71 a month in storage and under $4 in reads at 100,000 questions, so the Standard plan’s $50 monthly minimum is what you pay.15 Embedding the corpus once with OpenAI’s text-embedding-3-small, at $0.02 per million tokens, costs about $3.8
  • Observability. Langfuse’s Core plan costs $29 a month for 100,000 units, then $8 per 100,000, where a unit is a trace, an observation or a score.16
  • Evals. The RAG run is 300 questions, each answered and then scored by a judge call of the same size, for $12.60. The voice run is 100 simulated three-minute calls, with the simulated caller assumed to cost as much as the agent, for $44.70. The agent run is 200 tasks for $38.70. For the voice agent and the agentic workflow, eval runs cost more than hosting.

One-off and monthly costs for the three builds

This table puts the sections above together for 1,000 users at medium use, as of 8 October 2026. Upkeep is the labour that keeps the product working after launch. GoodFirms lists maintenance and support at 15% to 25% of an app budget, as a yearly cost,1 and the cost model applies that range to the build cost. Figures are rounded.

Build One-off build Monthly LLM spend Monthly hosting, vectors, evals, observability Monthly upkeep Monthly total
RAG assistant $31,700–$165,000 $105–$2,100 $398 $400–$3,400 $900–$5,900
Voice agent on a phone line $39,000–$202,500 $2,130 $1,047 $490–$4,200 $3,700–$7,400
Agentic workflow $47,500–$245,000 $488–$8,625 $617 $590–$5,100 $1,700–$14,300

Each range runs from payroll rates with the smaller model, assuming it passes your evals, to agency rates with a mid-tier model.

Over a year, running costs can overtake the build. The agentic workflow on Sonnet 5.5 at medium use spends $8,625 a month on tokens, or $103,500 a year, which is more than its whole build costs at $95 an hour.

Costs that first quotes often leave out

Each of these has a cost you can estimate before you sign.

  • Evals. Test cases with known good outputs, a harness that runs them on every change, and someone who reads the failures. In the cost model they take 40 to 160 hours to build and $126 to $447 a month to run. Without them, nobody can tell whether a cheaper model or a new prompt made the product worse.
  • Monitoring. Traces of every model call with its inputs, outputs, cost and latency, plus alerts when one of them drifts. The tooling is cheap, and most of the cost is the time to wire it in and to look at it.
  • Security review. The OWASP Top 10 for LLM applications (2025) starts with prompt injection, and also covers sensitive information disclosure, excessive agency, vector and embedding weaknesses and unbounded consumption.17 A RAG assistant that ignores document permissions, or an agent with write access to everything, fails review late, when fixes cost the most. GoodFirms puts security audits and legal compliance for an app at about $5,000 to $25,000 and up.1
  • Rate limits and spend controls. Provider rate limits cap throughput, so a job that stays under them can still spend a lot. We described how a backfill bug ran up a $1,100 vision bill, and the per-job ledgers and daily caps that catch it early.
  • Compliance. Article 50 of the EU AI Act has applied since 2 August 2026. Systems that talk to people must tell them they are dealing with an AI, unless that is obvious, and generative systems must mark their output in a machine-readable way. Generative systems placed on the EU market before 2 August have until 2 December 2026 for the marking, and fines can reach €15 million or 3% of worldwide turnover.18 Our post on what EU and California law require of AI-generated content maps the duties to features.
  • Model changes. Each new model or price change means another eval run, and sometimes prompt work.

Does AI-assisted coding make the build cheaper?

Some of the hours in the reference builds will go faster with coding agents. How much faster is less settled than many vendor pages suggest. GoodFirms’ guide says AI lets many businesses cut app development costs by 20% to 40%, without saying how that was measured.1

METR’s controlled studies are the most careful public measurements we know of. Its early-2025 study found that experienced open-source developers took 19% longer when AI tools were allowed. In February 2026 it reported a second round, with 57 developers and more than 800 tasks.19 Developers who returned from the first study took an estimated 18% less time with AI, with a confidence interval running from 38% less to 9% more. Newly recruited developers took 4% less, with an interval from 15% less to 9% more.

METR described its new data as “an unreliable signal of the current productivity effect of AI tools”. Between 30% and 50% of developers said they had held back tasks they didn’t want to do without AI, so METR thinks the true speedup is probably larger than it measured. It is redesigning the experiment.

The tools also cost money. On 24 June 2026, Gartner predicted that AI coding costs will overtake the average developer’s salary by 2028, as token use rises and vendors move to consumption-based pricing.20 It recommends sending simple tasks to smaller models and setting token thresholds.

Much of the effort in the reference builds goes on decisions and checks, such as which documents a user may see, what counts as a correct answer and which actions need a person. Faster code generation helps less with that work. For a budget, ask whoever quotes whether their hours already assume AI-assisted coding, and by how much, and put the coding tools’ token spend on the bill as its own line.

What this cost model assumes and leaves out

  • Every hour figure is an assumption built from the work packages above. None is a measurement from our own projects.
  • Rates and prices are as of 8 October 2026, from the pages cited. Model prices are standard pay-as-you-go rates, without batch or enterprise discounts.
  • The request shapes leave out reasoning tokens and cache-write fees, and the agent steps assume a warm cache.
  • Usage levels, eval sizes, run counts, corpus size and the hosting setup are assumptions. Change any of them and the arithmetic in each section still works.
  • It also leaves out product design beyond a basic interface, mobile apps, data labelling, fine-tuning, extra languages or regions, enterprise security questionnaires, legal review, taxes and growth past 1,000 users.

Questions to ask before you accept a quote

  1. Who does the work, and for how many hours each? Ask for each person’s seniority and hours, and whether any work is subcontracted. A blended rate hides the mix.
  2. What are the hours per work package? A quote you can compare has a breakdown like the table above. A fixed price without one still rests on an estimate of hours.
  3. Which of the forgotten costs are included? Check evals, observability, security review, deployment, documentation and handover by name.
  4. How will quality be measured before launch? Ask how many test cases the eval set will have, who writes them, what score counts as ready, and whether you keep the set.
  5. What will it cost to run? Ask for the model, the cost per request at your expected usage, and the spend limits and alerts. Ask whose account pays the model bills, and whether they are marked up.
  6. What happens when the model or its price changes? Find out who re-runs the evals and who pays for that time.
  7. Who owns the code, prompts, eval sets and data? They should live in your repository and your cloud accounts from the first week, with API keys in your name.
  8. How are changes priced? Ask what happens when the scope moves, and how much contingency the quote already carries.

  1. GoodFirms, “How Much Does It Cost to Develop an App in 2026? Full Development Cost Breakdown”, based on input from 267 mobile app development companies, updated 25 September 2026, https://www.goodfirms.co/resources/cost-to-develop-an-app ↩↩↩↩↩↩↩↩

  2. AxonBuild, “MVP Development Cost: What the Published Ladders Assume”, 15 September 2026, https://axonbuild.com/blog/mvp-development-cost ↩↩

  3. U.S. Bureau of Labor Statistics, “Software Developers, Quality Assurance Analysts, and Testers”, Occupational Outlook Handbook, pay data for May 2025, https://www.bls.gov/ooh/computer-and-information-technology/software-developers.htm ↩↩

  4. U.S. Bureau of Labor Statistics, “Employer Costs for Employee Compensation, June 2026”, 9 September 2026, https://www.bls.gov/news.release/ecec.nr0.htm, and Table 4, private industry workers by occupational group, https://www.bls.gov/news.release/ecec.t04.htm ↩↩

  5. Go Fractional, “Fractional Senior Software Engineer Cost & Rates”, updated 8 October 2026, https://www.gofractional.com/insights/rates/senior-software-engineer ↩

  6. Go Fractional, “Fractional CTO Cost & Rates”, updated 8 October 2026, https://www.gofractional.com/insights/rates/cto ↩

  7. Anthropic, “Pricing”, https://claude.com/pricing ↩

  8. OpenAI, “Pricing”, https://developers.openai.com/api/docs/pricing ↩↩↩

  9. OpenAI, “GPT-Live-1”, model page, https://developers.openai.com/api/docs/models/gpt-live-1 ↩

  10. Artificial Analysis, “LLM Leaderboard”, https://artificialanalysis.ai/leaderboards/models ↩

  11. Google, “Gemini Developer API pricing”, https://ai.google.dev/gemini-api/docs/pricing ↩

  12. OpenAI, “Cost optimization”, Realtime API guide, https://developers.openai.com/api/docs/guides/realtime-costs ↩

  13. Twilio, “Voice pricing: United States”, https://www.twilio.com/en-us/voice/pricing/us ↩

  14. Render, “Pricing”, https://render.com/pricing ↩

  15. Pinecone, “Pricing”, https://www.pinecone.io/pricing/, and “Understanding cost”, https://docs.pinecone.io/guides/manage-cost/understanding-cost ↩

  16. Langfuse, “Pricing”, https://langfuse.com/pricing ↩

  17. OWASP Gen AI Security Project, “Top 10 for LLM Applications 2025”, https://genai.owasp.org/llm-top-10/ ↩

  18. European Commission, “Transparency obligations under Article 50 of the AI Act”, questions and answers, https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act ↩

  19. METR, “We are Changing our Developer Productivity Experiment Design”, 24 February 2026, https://metr.org/blog/2026-02-24-uplift-update ↩

  20. Gartner, “Gartner Predicts AI Coding Costs Will Surpass Average Developer’s Salary by 2028 as Token Consumption Surges”, 24 June 2026, https://www.gartner.com/en/newsroom/press-releases/2026-06-24-gartner-predicts-ai-coding-costs-will-surpass-average-developer-salary-by-2028-as-token-consumption-surges ↩

Frequently asked questions

How much does it cost to build an AI agent?

On our cost model, a first production release takes about 330 to 980 engineering hours, or roughly $32,000 to $245,000 at $95 to $250 an hour, as of October 2026. An agent that takes actions in other systems sits at the top of that range because of integrations, approvals and testing.

What is the cost to build an MVP in 2026?

The MVP price ladders of six development companies, collected by AxonBuild in September 2026, have middle tiers that together span $15,000 to $100,000. In our cost model, AI MVP development cost sits near the low end of each range, about $32,000 to $48,000 at $95 an hour.

How much does an AI app cost to run per month?

For 1,000 users at moderate use, our cost model gives about $500 to $9,200 a month for model calls, hosting, vector storage, evals and observability, before upkeep labour. On a mid-tier model, tokens are the largest of those lines, and a smaller model can cut them by up to 20 times if it passes your evals.

Does AI-assisted coding make app development cheaper?

Possibly, but the best public measurement is weak. METR’s February 2026 update estimated that experienced developers took 4% to 18% less time with AI tools, with intervals that include a slowdown, and METR called the signal unreliable. Gartner expects AI coding costs to pass the average developer’s salary by 2028.

How much does an AI voice agent cost per minute?

As of 8 October 2026, OpenAI’s GPT-Live-1 costs $0.05 per minute of session, billed per second, plus normal rates for the backend model it calls. With one backend call a minute on GPT-6.1 Sol, our cost model gets $0.071 a minute, and a US phone line through Twilio adds about $0.013.

How much does a fractional CTO cost?

Go Fractional’s benchmark, updated 8 October 2026, puts the median at $200 an hour, with the middle half of rates between $150 and $250. At a typical 15 hours a week, that comes to $9,000 to $15,000 a month.

Work with us

Building something like this?

9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.