Get in touch

9io.ai / Blog

Guarding an LLM research pipeline with two reviews and a judge

How 9io Alpha checks every model answer with two differently prompted reviews, a judge, rule-based checks and a fallback for when the model fails.

Key takeaways

  • Check allowed values, ranges and required fields in code on every model answer.
  • Two prompts on one model share its errors, so their agreement is weaker evidence than it looks.
  • Keep a rules-based path that answers when the judge times out or returns something unusable.
  • Bound stated confidence with rules, such as a ceiling of 60 when the reviewers disagree.
  • Coalesce identical requests behind a per-key lock, and cache only results that succeeded.

Every verdict from 9io Alpha’s research tool goes through two reviews of the same evidence, a judge that makes the final call, and a layer of ordinary code that checks the judge’s answer before anyone sees it. If the judge times out or returns something unusable, fixed rules produce the verdict instead.

9io Alpha is our own trading research platform. This post walks through the pipeline behind its one-click research verdict. It covers how we gather evidence once per stock, why our two “desks” are less independent than they sound, which checks run on every answer, how the fallback decides, and how we bound time and token spend. It ends with what we’re changing and a checklist.

The post describes engineering and is not investment advice.

What runs when someone asks for a verdict

A request names one stock. Stage 0 builds an evidence pack for it. The pack holds a year of daily prices and the indicators computed from them, any setups found by our seven rule-based detectors, a set of longer-term factor scores, and recent social and news items. A call to a web-research API runs at the same time and returns a short summary of recent news, catalysts and risks.

Stage 1 asks for two reviews of that pack. We call them the momentum desk and the risk desk. One system prompt tells the model to find the best trade in the data. The other tells it to look for the reasons a trade would fail, and to vote DOWN or no edge unless the setup survives that scrutiny. Both desks are the same reasoning model reading the same evidence. A third vote comes from plain rules with no model involved. It votes UP when a setup detector fired, DOWN on a simple downtrend rule, and “no edge” otherwise.

Stage 2 sends the three votes and the key numbers from the pack to a judge, which is the same model again with a third prompt. The judge makes the final call and writes the thesis. Code checks that answer, and if the judge fails, a fixed merge of the three votes decides.

Stage What runs Time limit If it fails
0, evidence Prices, indicators, setup detectors, factor scores, social and news items None of its own; cached for 10 minutes Missing prices end the request with an error; the other parts are optional
0, web research One web-research call 45 s; cached for 1 hour The pipeline continues without it
1, desks Two model calls in parallel, plus the rules vote 75 s each A failed desk is dropped; the rules vote always exists
2, judge One model call 75 s A rule-based merge of the votes decides

One evidence fetch per stock, however many requests arrive

Several research commands read the same evidence pack, such as the verdict, a chart read and a free-text question. A person can run them back to back, and two people can ask about the same stock at once. Building the pack means a price download plus indicator and detector work, so we cache it for 10 minutes per market and stock.

A plain cache still has a gap. When an entry is missing or expired, every request that arrives before the first fetch finishes starts a fetch of its own. That pile-up is called a cache stampede, and one standard fix is to let a single caller recompute the value while the others wait1. Go’s x/sync package ships the same idea as singleflight, where duplicate callers wait for the call in flight and receive its result2. We do it with one asyncio lock per stock and a second cache check inside the lock.

import asyncio
import time

TTL_S = 600
_cache: dict[str, tuple[float, dict]] = {}
_locks: dict[str, asyncio.Lock] = {}

def _fresh(key: str) -> dict | None:
    hit = _cache.get(key)
    return hit[1] if hit and time.monotonic() - hit[0] < TTL_S else None

async def evidence(key: str, build) -> dict:
    if (pack := _fresh(key)) is not None:
        return pack
    async with _locks.setdefault(key, asyncio.Lock()):
        if (pack := _fresh(key)) is not None:  # filled while we waited
            return pack
        pack = await build(key)                # raises on failure, so nothing is cached
        _cache[key] = (time.monotonic(), pack)
        return pack

The second check is what turns waiting callers into cache hits. The other detail is failure. If the fetch raises, nothing is cached, and the next waiter takes the lock and tries again. Singleflight instead hands the same error to every waiter. Retrying in turn suits a price source that fails briefly and recovers. During a long outage it costs time, because each waiter has to see its own attempt fail.

An asyncio lock only coordinates requests inside one process. With several worker processes, each has its own cache and its own locks, and you would move both into a shared store.

Any other cached call needs the same two rules, one lock per key and a cache that keeps only replies that parsed. A cached failure is worse than no cache, because it repeats the failure for the cache’s whole lifetime.

Two desks from one model

The two desks are a deliberate argument. One prompt leans toward action and the other toward doubt, and the judge reads both. The idea comes from two lines of work. Sampling several reasoning paths and keeping the most consistent answer improves on a single greedy pass3. Letting several model instances debate before answering improves reasoning and factual accuracy4.

The desks are less independent than two separate reviewers would be. Both are the same model reading the same evidence, and only the system prompt changes. The errors of different language models are often correlated, even across vendors. A study of more than 350 models found that on one leaderboard, when two models both got a question wrong, they gave the same wrong answer 60% of the time. Shared providers and architectures raised the correlation further5. Two prompts on one model sit at the far end of that scale, so when both desks agree, part of the agreement is the model agreeing with itself.

Research on mixing models points the same way. ReConcile, a method in which different models discuss a question and vote, reports that the diversity from using different models is critical to its results6. A benchmark of multi-agent debate methods found they don’t reliably beat simpler strategies such as self-consistency and ensembling7. We now treat the desks as two framings of one opinion. That is useful for making the bear case explicit, and it is weaker evidence than two different models agreeing.

What we would change, cheapest first:

  • Sample each desk twice and record whether it agrees with itself. A desk that flips between runs is giving an unstable answer.
  • Give the desks different slices of the evidence, so they can’t both anchor on the same few numbers.
  • Move one desk to a different model family. The code already has a path for a second provider, which is switched off today.

What the judge sees, and the biases it may bring

The judge receives each vote with its levels, thesis, reasons and risks. It also gets key numbers from the pack, such as the price, the 20-day and 52-week ranges and the ATR. ATR is average true range, a volatility measure usually taken over 14 days8. The judge returns a verdict, a confidence from 0 to 100, an action from a fixed list, entry, target and stop levels, and a thesis of four to six sentences.

LLM judges have documented biases. Zheng and colleagues found that strong models agree with human preferences in more than 80% of chat comparisons, about as often as humans agree with each other. They also documented position, verbosity and self-enhancement biases9. Wang and colleagues showed that changing the order in which answers appear can flip a judge’s ranking. With a large commercial model as the judge, reordering alone let a 13-billion-parameter open model beat that commercial model on 66 of 80 queries10.

Our judge always sees the momentum desk first, then the risk desk, then the rules vote. We haven’t tested whether swapping the order changes its verdict. The test is cheap. Rerun a sample of requests with the desks swapped and log every verdict that changes. If many do, Wang and colleagues’ calibration, which combines results across orders, is the next step.

The judge’s confidence is a number the model wrote. Models tend to be overconfident when they state their confidence in words11, so we treat it as an input for rules to bound. The cap that applies when the desks disagree is one of those rules.

Checks that run on every answer

JSON mode makes a model return parseable JSON. It promises nothing about which fields are present or what they contain. Two major API vendors document schema-constrained output as the stronger option, and both note that it can still break when the model refuses or runs out of tokens1213. Our live calls use JSON mode, so code does all of the checking.

Check Rule On failure
Parse Strip code fences and a leading “json” tag, then parse The answer counts as missing
Allowed values Verdict must be UP, DOWN or NO_EDGE after upper-casing The desk vote is dropped, or the judge answer rejected
Action Must come from a fixed list Derived from the verdict
Ranges Confidence 0 to 100; horizon 1 to 5 days Clamped
Rationale The judge must write a thesis The judge answer is rejected
Disagreement The two desks made different calls Confidence capped at 60
Levels Target within 8% and 3.5 × ATR of entry; stop within 6% Clamped
Reward-to-risk A long at market below 1.2 Relabelled “buy on pullback”, with a note on the level to wait for

Upper-casing first matters even with schemas, because one vendor notes that constrained output can return an enum value that differs only in capitalisation13. The parser strips code fences because the web-research call has no JSON mode, and its replies can arrive wrapped in a Markdown block.

Price levels are clamped instead of rejected, because throwing away a sound direction over one ambitious number would send more requests to the fallback. A short version of the clamp for a long:

def clamp_long(entry: float, target: float, stop: float, atr: float):
    """Bound a long's levels for a 1-5 day hold. Returns (target, stop, reward_to_risk)."""
    target = min(target, entry * 1.08, entry + 3.5 * atr)  # at most 8% and 3.5 ATR away
    target = max(target, entry * 1.005)                    # the caps must not cross the entry
    stop = min(max(stop, entry * 0.94), entry * 0.995)     # within 6%, and below the entry
    return target, stop, (target - entry) / (entry - stop)

def action_for_long(action: str, reward_to_risk: float) -> str:
    return "BUY_PULLBACK" if action == "BUY_NOW" and reward_to_risk < 1.2 else action

A clamp that works relative to the model’s own entry can’t catch an entry that is far from the market, so the entry needs its own bound, such as a band of a few ATR around the last price. These levels can feed an order ticket, where the order code runs its own checks. We cover those in order safety by default in code that can place trades.

When the judge fails, rules decide

The judge fails in ordinary ways. It times out, the API returns an error, the reply doesn’t parse, or it parses without a thesis. In each case the code falls back to a fixed merge of the votes it already has.

  1. If both desks answered and agree, their call wins. Confidence is their average plus 10, capped at 95.
  2. If they split, the rules vote wins when its confidence is at least 65, and the final confidence is capped at 60.
  3. Otherwise the verdict is “no edge”.

With one desk, its call stands with 10 points less confidence, restored if the rules vote agrees. With no model key configured, no model call is made and the rules vote becomes the verdict. This merge is what the pipeline used before the judge existed, so a judge failure drops back to the previous behaviour instead of to an error.

The merge has an asymmetry. The rules vote’s detectors only look for long setups, and its single downtrend rule carries a fixed confidence of 50, below the 65 needed to settle a split. So when the desks disagree, the fallback can resolve to UP or to no edge and never to DOWN.

Each response records which path produced the verdict, how many desk votes arrived and whether web research was available. We log that but don’t aggregate it yet, so we can’t say how often the fallback runs. That rate, and how often the desks disagree, are the first two numbers we would chart.

Keeping time and token spend bounded

Each stage runs under a hard time limit of 45 seconds for web research, 75 for each desk and 75 for the judge. The desks run in parallel, and web research overlaps the evidence fetch. Stacked end to end, the limits still allow 195 seconds, which is a long time to watch a spinner. The evidence fetch sits outside these limits and has no limit of its own, which we should add.

Every model call is capped at 8,000 output tokens at medium reasoning effort. On both of those vendors’ APIs, reasoning tokens count against the output cap1415. One vendor suggests reserving at least 25,000 tokens for reasoning and output when you start out14. We chose the lower cap because the answers are short JSON objects. The cost of that choice is that a long chain of reasoning can use up the budget before any answer is written. The parser then sees an empty or truncated reply, and the fallback runs. Across the three model calls, the caps bound a request at 24,000 output tokens plus its input.

Caching and a few hard limits keep spend down. The caches mean a busy stock costs one evidence fetch every 10 minutes and roughly one research call an hour. No model or research call is made when no key is configured. The batch endpoints that run model analysis over a list of stocks reject lists longer than 10 stocks for a deep scan and 20 for a quick scan. The deep scan analyses its stocks concurrently, so that cap is also its only concurrency limit.

What we’re changing

Writing this post meant reading the code against its own comments and prompts, and four changes came out of it.

  • Labels that match their sources. The prompts called each setup’s figure a “measured win rate”, when in this path it is a hard-coded starting assumption scaled by setup quality. A model told that a number was measured will treat it as evidence, and the judge is asked to cite such figures in its thesis, so the label reaches the text a person reads. Every number in a prompt now needs to say whether it was measured or assumed.
  • Schemas on the live path. JSON mode plus code checks has held up, but sending the schema as well costs nothing and catches malformed answers earlier.
  • More independent desks. Different slices of evidence per desk first, then a second model family.
  • An order test for the judge. Rerun a sample with the desks swapped, and log every verdict that changes.

A checklist for an LLM pipeline that makes decisions

  • Check allowed values, ranges and required fields in code on every answer, even with schema-constrained output.
  • Bound any number that flows into an action, such as a price, size or date, instead of rejecting the whole answer.
  • Cap stated confidence with rules you can explain, like a ceiling when the reviewers disagree.
  • Keep a deterministic path that answers when the model can’t, and record which path answered.
  • Use different model families when you want independent reviews.
  • Test a judge with its inputs in a different order before trusting its tie-breaks.
  • Coalesce identical requests in flight, and cache only results that succeeded.
  • Give every stage a time limit, and every call an output cap that counts reasoning tokens.
  • Tag each number you put in a prompt as measured or assumed.

  1. Cache stampede, Wikipedia, https://en.wikipedia.org/wiki/Cache_stampede ↩

  2. Package singleflight, golang.org/x/sync, https://pkg.go.dev/golang.org/x/sync/singleflight ↩

  3. Wang et al., “Self-Consistency Improves Chain of Thought Reasoning in Language Models”, ICLR 2023, https://arxiv.org/abs/2203.11171 ↩

  4. Du et al., “Improving Factuality and Reasoning in Language Models through Multiagent Debate”, 2023, https://arxiv.org/abs/2305.14325 ↩

  5. Kim, Garg, Peng and Garg, “Correlated Errors in Large Language Models”, ICML 2025, https://arxiv.org/abs/2506.07962 ↩

  6. Chen et al., “ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs”, ACL 2024, https://arxiv.org/abs/2309.13007 ↩

  7. Smit et al., “Should we be going MAD? A Look at Multi-Agent Debate Strategies for LLMs”, 2023, https://arxiv.org/abs/2311.17371 ↩

  8. Average true range, Wikipedia, https://en.wikipedia.org/wiki/Average_true_range ↩

  9. Zheng et al., “Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”, NeurIPS 2023 Datasets and Benchmarks Track, https://arxiv.org/abs/2306.05685 ↩

  10. Wang et al., “Large Language Models are not Fair Evaluators”, 2023, https://arxiv.org/abs/2305.17926 ↩

  11. Xiong et al., “Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMs”, ICLR 2024, https://arxiv.org/abs/2306.13063 ↩

  12. OpenAI, Structured Outputs guide, https://developers.openai.com/api/docs/guides/structured-outputs ↩

  13. Anthropic, Structured outputs, https://platform.claude.com/docs/en/build-with-claude/structured-outputs ↩↩

  14. OpenAI, Reasoning models guide, https://developers.openai.com/api/docs/guides/reasoning ↩↩

  15. Anthropic, Extended thinking, https://platform.claude.com/docs/en/build-with-claude/extended-thinking ↩

Frequently asked questions

What is an LLM-as-a-judge pipeline?

Several model calls review the same input, and a further call reads their answers and makes the final decision. In a production system, code should then check that decision before anyone sees it.

Are two prompts on the same model independent reviewers?

No. They share the model’s training and blind spots, so their errors are correlated. A study of more than 350 models found that errors correlate more when models share a provider or architecture.

Is JSON mode enough to get reliable structured output from an LLM?

JSON mode only promises parseable JSON. Schema-constrained output also enforces fields and types, but vendors document that it can still break on refusals or token limits, so validate in code as well.

What should happen when the judge model fails?

Answer from deterministic rules and record that the fallback ran. In our pipeline, agreeing reviewers win, a confident rules-based vote settles a split, and otherwise the verdict is no edge.

How do you keep an LLM pipeline’s latency and cost bounded?

Give each stage a hard time limit, cap output tokens with reasoning tokens counted, cache shared inputs, skip model calls when no key is configured, and cap batch sizes.

Work with us

Building something like this?

9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.