Key takeaways
- Distinct topic pairs scored 0.67 to 0.91 on cosine similarity; reworded duplicates scored 0.70 to 0.94.
- Ranges that overlap that much leave no threshold that separates duplicates from neighbouring topics.
- Merge decisions now use two rules, one on learning-objective verbs and one on near-identical sentences.
- A rare-word guard refused 16 of 18 true duplicates, so we deleted it.
- Make each outline topic cite a page range and a quote that must appear on that page.
Cosine similarity can’t decide on its own whether two topics are the same topic. We found this while building an engine that turns textbooks into courses. With OpenAI’s text-embedding-3-large, pairs of topics that were similar but distinct scored between 0.67 and 0.91. Pairs that were rewordings of the same topic scored between 0.70 and 0.94. The ranges overlap from 0.70 to 0.91, so no threshold merges the duplicates without also merging topics that should stay apart.
The engine is part of an AI study platform for high-school and college students that we build and run. It takes a set of textbook PDFs and produces an interactive course. The course is an outline of sections, and each section has explanations, diagrams, simulations, flashcards and quizzes. When two books cover the same topic, the outline should have one section for it.
This post walks through the pipeline, a bug that changed how it reads books, the check that ties the outline to the pages, why the threshold failed and the rules that replaced it. The measurements, how we took them, and a problem we haven’t fixed yet are near the end.
How a set of PDFs becomes a course
A build fans out across a fleet of Python workers. Each module of the course goes onto an Amazon SQS queue as its own message, and whichever worker is free picks it up. The rest of the stack is in the table.
| Part | What we use | What it does here |
|---|---|---|
| Queue | Amazon SQS, one message per module | Spreads a build across the workers |
| Workers | Python processes | Run each module’s steps |
| Rate limiting | A Redis-backed limiter shared by the whole fleet | Gives every worker one shared budget for model calls |
| Storage | One MongoDB database per tenant | Keeps each tenant’s data in its own database |
| Text extraction | PyMuPDF | Reads the text of every page |
| Embeddings | OpenAI text-embedding-3-large | Compares topics across books |
The limiter lives in Redis because a limit enforced inside one process can’t see the other processes, and a model provider’s limits apply to all of them together. Redis documents the basic building block, a counter per time window made from INCR and EXPIRE 1. Throttling in this pipeline has a story of its own, told in our post on a pipeline that reported 0 dropped while most of its output was missing.
A database per tenant is one of the layouts MongoDB describes for multi-tenant systems. Its guide lists the trade-offs. Access can be restricted per database, and one tenant can be backed up, restored or moved to its own cluster without touching the rest. The cost is that every tenant repeats the same collections and indexes, and each of those uses resources on the cluster 2.
The first version read part of the books
The first version failed quietly. Given several PDFs, it read only the first 40 pages, or built the course from one of the books and ignored the others. Nothing raised an error. A language model writes a fluent outline from whatever text it receives, so a course built from 40 pages can look complete.
We changed two things. The engine now reads every page of every book. And it builds a fixed outline of topics from all the books up front.
Reading every page costs little. In our runs PyMuPDF read 244 pages in 0.98 s and 883 pages in 3.8 s, about 4 ms a page. The project publishes its own comparison, in which PyMuPDF extracts text 3.4 times faster than XPDF and 28 times faster than PDFMiner on the project’s test set 3. That comparison is vendor-run, so treat it as a rough guide.
An outline that has to quote its pages
Each topic in the outline claims a page range and carries a quote from the book. We check each topic against the page range it claims, and its quote has to actually appear on that page.
This catches a specific failure. A model can write a believable topic list for a subject without reading the book in front of it. It is very unlikely to produce a sentence printed on page 212 of that book unless the page was in its input. String matching is also cheap and deterministic, so there is no threshold to tune.
PDF text needs normalising before you match against it. Ligatures can come out as single characters, so “fi” arrives as “fi”. Words are split with a hyphen at the end of a line. Quote marks may be curly in the PDF and straight in the model’s output. PyMuPDF has extraction flags for the first two. Clearing TEXT_PRESERVE_LIGATURES expands ligatures, and TEXT_DEHYPHENATE joins a word that was hyphenated across a line break 4. Its search_for method ignores ASCII case and detects line-end hyphens by default 5, which covers many cases. If you’d rather own the normalisation, the check can look like this (written for this post):
import re
import pymupdf
# Expand ligatures and join words hyphenated across line breaks.
FLAGS = (pymupdf.TEXTFLAGS_TEXT & ~pymupdf.TEXT_PRESERVE_LIGATURES) | pymupdf.TEXT_DEHYPHENATE
# Straighten curly quotes and drop hyphens, so "well-known" matches "wellknown".
MARKS = str.maketrans({"\u2018": "'", "\u2019": "'", "\u201c": '"', "\u201d": '"', "-": None})
def norm(text: str) -> str:
return re.sub(r"\s+", " ", text.translate(MARKS)).strip().casefold()
def quote_on_pages(doc: pymupdf.Document, quote: str, first: int, last: int) -> bool:
"""True if the quote appears on a page in first..last (1-based, inclusive)."""
needle = norm(quote)
return any(needle in norm(doc[i].get_text("text", flags=FLAGS))
for i in range(first - 1, last))
Why cosine similarity could not decide merges
Several books on one subject cover many of the same topics, each in its own words. One build folded 280 candidate topics into 77 sections, and every merge is a decision. The engine has to tell two books’ treatments of one topic, which become one section, from neighbouring topics with similar wording, which stay separate.
The obvious tool is embedding similarity. You embed each topic, compute cosine similarity for candidate pairs and merge the pairs above a threshold. OpenAI recommends cosine similarity for its embeddings. The vectors are normalised to length 1, so cosine similarity and Euclidean distance put pairs in the same order 6. Switching the distance measure changes nothing, which leaves the threshold as the only choice.
We measured both kinds of pair.
| Pair | Cosine similarity (text-embedding-3-large) |
|---|---|
| Similar but distinct topics | 0.67 to 0.91 |
| Rewordings of the same topic | 0.70 to 0.94 |
Each range is 0.24 wide, and 0.21 of that is shared. A threshold of 0.92 would have merged none of the distinct pairs and refused every duplicate that scored below 0.92. A threshold of 0.70 would have merged every duplicate, and every distinct pair from 0.70 to 0.91 along with them. Any threshold in between makes both mistakes.
Three things push the two kinds of pair together.
Embeddings are trained for relatedness. OpenAI hasn’t published how text-embedding-3-large was trained. Its paper on an earlier generation of embedding models describes training on neighbouring pieces of internet text as positive pairs. The paper also suggests that search relevance and sentence similarity can pull against each other, since a sentence and its negation are relevant to one another without meaning the same thing 7. Two sibling topics from one chapter are about as related as two pieces of text get.
Scores bunch together. Ethayarajh found that the contextual word representations in ELMo, BERT and GPT-2 are anisotropic. They sit in a narrow cone instead of spreading through the space, and in layers 2 to 8 of GPT-2 two random words had an average cosine similarity of roughly 0.6 8. Li and colleagues found the same problem in BERT sentence embeddings and showed that it hurts similarity scoring 9. We didn’t measure anisotropy in text-embedding-3-large. Every related pair we compared, duplicate or not, landed between 0.67 and 0.94.
Small edits that change the meaning barely move the score. PAWS is a paraphrase dataset built from sentence pairs with heavy word overlap, and models trained on earlier paraphrase datasets scored under 40% accuracy on it 10. NevIR asked retrieval models to rank documents that differ only by a negation. Most did the same as or worse than random ranking, and bi-encoders, the architecture behind embedding search, were among the weakest 11. In a course outline, “Newton’s first law” and “Newton’s second law” differ by one word, and that word is the entire difference between the topics. That pair, like the other topic examples in this post, is invented.
Two rules that replaced the threshold
The merge decision now uses two rules.
The first is a check on the verb of each topic’s learning objective. In the revised Bloom’s taxonomy, an objective pairs a noun phrase, the content, with a verb phrase, the cognitive process the student carries out 12. Two objectives can share their content and differ in the verb, as “state Ohm’s law” and “use Ohm’s law to find the current in a circuit” do. The verb check asks whether two objectives want the student to do the same kind of thing with the material.
The second rule is a detector for sentences that differ by only one or two words. Those are the pairs where the score tells you least, because most of each sentence is shared. “Stages of mitosis” and “phases of mitosis” are the same topic, while “Newton’s first law” and “Newton’s second law” are two topics, and each pair differs by a single word. The detector flags these pairs so that a rule decides them instead of the score. Finding the changed words takes a few lines of standard-library Python (written for this post):
from difflib import SequenceMatcher
def changed_words(a: str, b: str) -> list[tuple[str, str]]:
"""Word-level edits between two short sentences, as (before, after) pairs."""
wa, wb = a.lower().split(), b.lower().split()
ops = SequenceMatcher(a=wa, b=wb, autojunk=False).get_opcodes()
return [(" ".join(wa[i1:i2]), " ".join(wb[j1:j2]))
for tag, i1, i2, j1, j2 in ops if tag != "equal"]
changed_words("Explain Newton's first law of motion",
"Explain Newton's second law of motion")
# [('first', 'second')]
Before these rules we had a guard that refused a merge when one side contained a rare word the other lacked. A guard like that assumes a rare word marks a specific concept. On true duplicates it refused 16 of 18, which means 16 of those 18 rewordings had a rare word on one side only. Two authors describing the same topic seldom choose the same uncommon words. We deleted the guard.
What the rules did on real pairs
| Approach | Pairs | Outcome |
|---|---|---|
| Cosine threshold | Distinct pairs and rewordings (ranges above) | No threshold separates them |
| Rare-word guard, now deleted | 18 true duplicates | Refused 16 |
| Verb check and one-or-two-word detector | A set of 13 real pairs | 0 wrong merges |
| Verb check and one-or-two-word detector | A set of 25 real pairs | 1 refusal |
On the 13-pair set the rules made no wrong merges, and on the 25-pair set they refused one merge. We give the results as counts. A percentage computed on 13 or 25 pairs would suggest more precision than sets this small can carry.
How we measured
- Similarity scores. Cosine similarity between the text-embedding-3-large embeddings of the two topics in each pair. We report each group as its minimum and maximum, because the extremes decide whether any threshold can separate the groups.
- Rule results. The 13-pair and 25-pair sets are real topic pairs. They are small, so they show how the rules behave on cases we know, and they can’t support an error rate.
- Rare-word guard. Run against 18 true duplicates before we removed it.
- Extraction. The time PyMuPDF took to read the text of two inputs, one of 244 pages and one of 883. This covers the text layer only. Scanned pages without a text layer need OCR, which PyMuPDF’s documentation puts at about a thousand times slower than normal extraction 13.
- Merging at build scale. The 280 candidate topics and 77 sections come from one build.
- Not covered here. This post doesn’t compare the rules with a cross-encoder, or with an LLM asked to judge each pair. NevIR’s results suggest that cross-encoders, which read both texts together, handle small changes in meaning better than bi-encoders do 11.
Still open: two workers on one module
Occasionally two workers pick up the same module. SQS delivers messages at least once. A message that hasn’t been deleted becomes visible again when its visibility timeout runs out (30 s by default), and AWS’s documentation says there is no absolute guarantee against a second delivery even inside the timeout 14. Extending the timeout while a worker runs, with ChangeMessageVisibility as a heartbeat, narrows the window. It can’t close it, because a worker can stall after its lease lapses and then write anyway.
The fix we’ve designed uses fencing tokens, as described in Martin Kleppmann’s article on distributed locking 15. Each time a worker claims a module, it gets a number higher than any issued before. The store remembers the highest number it has accepted and rejects writes that carry a lower one, so a stalled worker’s late write is refused. It isn’t built yet.
Here is the idea in MongoDB terms, as a sketch written for this post. The database that holds the result also issues the token, so there’s no second system to keep in step. MongoDB applies an update to a single document atomically, and putting the expected value in the update’s filter is its documented way to stop concurrent writers from overwriting each other 16.
from pymongo import ReturnDocument
def claim(coll, module_id) -> int:
"""Take the next fencing token for a module."""
doc = coll.find_one_and_update(
{"_id": module_id},
{"$inc": {"fence": 1}},
upsert=True,
return_document=ReturnDocument.AFTER,
)
return doc["fence"]
def save(coll, module_id, token: int, result: dict) -> bool:
"""Write only if no one has claimed the module since. False means a newer worker owns it."""
res = coll.update_one({"_id": module_id, "fence": token}, {"$set": {"result": result}})
return res.matched_count == 1
If a module writes many documents, every one of those writes needs the same check, for example inside a transaction.
A checklist before you trust a similarity threshold
- Label a few dozen pairs from your own data, true duplicates and close neighbours, and look at both score ranges before you pick a threshold. If the ranges overlap, no threshold will separate them.
- Use similarity to find candidate pairs. Decide merges with something that looks at what differs between the two texts.
- Give pairs that differ by one or two words their own handling. The score says least about exactly those pairs.
- Test every guard against true duplicates as well as distinct pairs. Our rare-word guard refused 16 of 18 duplicates.
- Make each generated topic cite a page range and a verbatim quote, and check both against the extracted text.
- Normalise ligatures, line-end hyphens and quote marks before matching quotes against PDF text.
- Decide whether a page number means the PDF’s page index or the number printed on the page. Front matter makes the two differ.
- Report the pages that no topic claims. A course built from the first 40 pages would leave most of its source unclaimed.
- If your queue can deliver a message twice, make the store reject stale writes. Fencing tokens do it with one conditional update.
-
Redis documentation, INCR command, “Pattern: rate limiter”, https://redis.io/docs/latest/commands/incr/. ↩
-
MongoDB Atlas documentation, “Build a Multi-Tenant Architecture”, https://www.mongodb.com/docs/atlas/build-multi-tenant-arch/. ↩
-
PyMuPDF documentation, “Performance Comparison Methodology”, https://pymupdf.readthedocs.io/en/latest/app4.html. ↩
-
PyMuPDF documentation, “Constants and Enumerations: Text Extraction Flags”, https://pymupdf.readthedocs.io/en/latest/vars.html. ↩
-
PyMuPDF documentation, Page.search_for, https://pymupdf.readthedocs.io/en/latest/page.html#Page.search_for. ↩
-
OpenAI API documentation, Embeddings guide, https://developers.openai.com/api/docs/guides/embeddings. ↩
-
Neelakantan et al., “Text and Code Embeddings by Contrastive Pre-Training”, 2022, https://arxiv.org/abs/2201.10005. ↩
-
Kawin Ethayarajh, “How Contextual are Contextualized Word Representations?”, EMNLP 2019, https://aclanthology.org/D19-1006/. ↩
-
Li et al., “On the Sentence Embeddings from Pre-trained Language Models”, EMNLP 2020, https://aclanthology.org/2020.emnlp-main.733/. ↩
-
Zhang, Baldridge and He, “PAWS: Paraphrase Adversaries from Word Scrambling”, NAACL 2019, https://aclanthology.org/N19-1131/. ↩
-
Weller, Lawrie and Van Durme, “NevIR: Negation in Neural Information Retrieval”, EACL 2024, https://arxiv.org/abs/2305.07614. ↩↩
-
David R. Krathwohl, “A Revision of Bloom’s Taxonomy: An Overview”, Theory Into Practice 41(4), 2002, https://doi.org/10.1207/s15430421tip4104_2. ↩
-
PyMuPDF documentation, “OCR: Optical Character Recognition”, https://pymupdf.readthedocs.io/en/latest/recipes-ocr.html. ↩
-
Amazon SQS Developer Guide, “Amazon SQS visibility timeout”, https://docs.aws.amazon.com/AWSSimpleQueueService/latest/SQSDeveloperGuide/sqs-visibility-timeout.html. ↩
-
Martin Kleppmann, “How to do distributed locking”, 8 February 2016, https://martin.kleppmann.com/2016/02/08/how-to-do-distributed-locking.html. ↩
-
MongoDB documentation, “Atomicity and Transactions”, https://www.mongodb.com/docs/manual/core/write-operations-atomicity/. ↩
Frequently asked questions
What cosine similarity threshold should I use to detect duplicate topics?
Measure before you pick one. On our textbook topics, distinct pairs scored 0.67 to 0.91 and reworded duplicates 0.70 to 0.94, so every threshold either merged distinct topics or missed duplicates. If your two ranges overlap like that, decide merges with something other than the score.
Why do different topics from the same subject get such high cosine similarity?
Embedding models are trained to place related text close together, and sibling topics in one subject are closely related. Studies of earlier language models also found that their vectors bunch into a narrow region, which compresses the range of scores.
Does switching from cosine similarity to Euclidean distance help?
Not with OpenAI embeddings. They are normalised to length 1, so cosine similarity and Euclidean distance rank pairs in the same order, and a threshold on one maps to a threshold on the other.
How do you check that an LLM-generated outline is grounded in the source PDF?
Have each topic name the page range it covers and a quote from that page, then look for the quote in the extracted text of that page. Normalise ligatures, line-end hyphens and quote marks before matching.
How fast is PyMuPDF at extracting text from a textbook?
In our runs it read 244 pages in 0.98 s and 883 pages in 3.8 s, about 4 ms a page. Scanned pages without a text layer need OCR, which PyMuPDF’s documentation puts at about a thousand times slower.
What is a fencing token?
A number that goes up every time a worker is granted a lock or lease. The storage layer rejects writes that carry an older number, so a worker whose lease expired can’t overwrite newer work.
Building something like this?
9io is a small team of senior engineers with a fractional CTO, and we work by the hour. Send us a note about your product. The reply comes from the person who'd do the work.