Get in touch
Blog

Engineering notes from 9io.

How we build products, AI systems and trading infrastructure — and what we learn on the way. Written by the people who do the work, with the sources linked.

How much does it cost to build an AI agent or app in 2026?

How much does it cost to build an AI agent? A cost model that shows its working, from build hours and market rates to monthly running costs per 1,000 users.

8 October 2026 · 16 min read

AI agent containment lessons from the summer 2026 incidents

Lab agents reached real systems and MCP tools turned on their users. What failed in each 2026 incident, and the controls to enforce outside the model.

8 October 2026 · 12 min read

AI decision models like Jev, and when to use one instead of an LLM

What an AI decision model like Jev is, how it differs from an LLM, the options as of 7 October 2026, what independent tests found, and when to use each.

7 October 2026 · 17 min read

Comparing FashionSigLIP, CLIP and DINOv2 for visual product search

On 60 real cases, FashionSigLIP put the right product in the top five 63% of the time, against 38% for CLIP and 23% for DINOv2. The pipeline and the test.

6 October 2026 · 9 min read

Migrating MCP servers and clients to the stateless 2026-07-28 spec

MCP’s 2026-07-28 revision drops sessions and the initialize handshake. What breaks, which clients support it in October 2026, and how to migrate safely.

6 October 2026 · 13 min read

How we halved an AI stylist’s latency, and when hedging stops paying

A streaming AI stylist went from a 30 s median response to 13 to 15 s. The four changes, how we measured them, and why hedged requests stopped paying.

6 October 2026 · 11 min read

Claude Code cost per month, per engineer, and how to cap the bill

Claude Code cost per month per engineer, Codex and Cursor plans compared, which spending caps are hard, and a policy template. As of 5 October 2026.

5 October 2026 · 18 min read

Is Claude nerfed? How to measure model drift yourself

Is Claude nerfed, or did something else change? Documented degradations, the changes that only look like one, and a nightly drift check with error bars.

3 October 2026 · 17 min read

How a backfill bug ran up a $1,100 LLM vision bill

A job re-classified items it had already classified: 419,057 vision calls and about $1,100. What went wrong, and the guardrails that catch it early.

2 October 2026 · 10 min read

How AI image pipelines fail quietly, and how to catch it

A Gemini try-on that took nearly two minutes to fail, product cutouts saved onto black, and the retry, cache and alpha fixes for each.

2 October 2026 · 11 min read

Grounding a study assistant in live web search and the course itself

A web mode that never searched, an SDK that turned 429s into AttributeErrors, and the numbers from answering with a whole course in the context.

1 October 2026 · 11 min read

What we learned building a voice tutor that draws while it talks

What broke in three versions of a voice tutor that draws as it talks, built on OpenAI’s Realtime API and GPT-Live, and the API details behind each problem.

29 September 2026 · 11 min read

How we cut an AI assistant’s time to first word to 3 seconds

An in-app AI assistant took 20 to 30 seconds to start answering. Putting its context in the prompt and forcing fewer tool calls got it to about 3 seconds.

29 September 2026 · 10 min read

Making one MCP server work in Claude, ChatGPT, Cursor and VS Code

What broke when one remote MCP server met Claude, ChatGPT, Cursor, VS Code and command-line clients, and the OAuth and version fixes behind each.

26 September 2026 · 13 min read

Guarding an LLM research pipeline with two reviews and a judge

How 9io Alpha checks every model answer with two differently prompted reviews, a judge, rule-based checks and a fallback for when the model fails.

15 September 2026 · 13 min read

Why cosine similarity couldn’t decide which textbook topics to merge

Distinct topics scored 0.67 to 0.91 and reworded duplicates 0.70 to 0.94, so no threshold worked. How we merge topics when turning textbooks into courses.

1 September 2026 · 12 min read

Why our LLM pipeline reported 0 dropped while most output was missing

A course build shipped 26 of 51 diagrams and reported 0 dropped. How a per-call limiter, 429s and deadlines hid the gap, where tokens went, and the fixes.

1 September 2026 · 11 min read