And past a gate that keeps unqualified questions out of the production bank.
Upstream is turning thousands of past-paper PDFs into a bank that is retrievable, classifiable and QA-able question by question. Downstream is collapsing one natural-language sentence into a set of executable retrieval constraints. This page is how both were built —— and what I fell into along the way.

No microservices at MVP stage. The selection criterion was exactly one: can a single person deploy and roll this back reliably?
So the MVP could close the loop first and swap in real implementations one at a time:
| AI | mock / openai-compatiblepaper building is wired to DeepSeek; homework/report generation is still mock |
| Storage | local / cos / wechat-cloudthe last two are placeholders, no real SDK yet |
| Repository | PrismaRepository / MemoryRepositorytests don't need a database |
The easy way to build natural-language paper generation is to throw the whole sentence at an LLM and ask for JSON back. I didn't, because that means every single turn costs money, adds latency, and can be unstable. What ships is a four-level funnel: whatever an earlier level can catch never reaches the next.
Topic lookup supports bare topics —— a teacher says "integration" without naming a unit, and the system still has to find it. Which means retrieval must be able to resolve globally, with no unit context.
And once global resolution is allowed, the P0 below becomes inevitable.
Intermittent, no pattern, and the tests asserted "pass". The front end showed nothing wrong either —— the unresolved-slots list was empty, so the UI looked perfectly healthy.
U1/U2/U4)subject was dropped while assembling parametersAssume "higher number means harder". Two problems: the assumption was never validated, and the same request generated twice produced a drifting difficulty spread. When a teacher asked "on what basis is this one hard", there was no answer.
Difficulty is the question's percentile inside its own unit. Same request twice, same difficulty spread; and every band is explainable as "this band means this percentile range within the unit".
Absolute difficulty isn't comparable across A-Level units in the first place —— "hard" in P1 and "hard" in FP3 are not the same thing. When a teacher says "P1, a bit harder", they mean harder within P1, not on some cross-unit absolute scale. A percentile expresses exactly that.
The side effect: newly loaded questions have no difficulty value until that unit's distribution is recomputed. That is the direct cause of the "difficulty not set" above, and the price this design charges.
Whether paper building is usable at all comes down to one tedious thing —— whether every question was filed under the right topic. Having a teacher review thousands one by one isn't realistic; handing it entirely to a model isn't trustworthy.
High confidence, both agents agree → straight into the production bank, retrievable.
Disagreement or middling confidence → into the review queue for a teacher to settle. Never ships to production.
Low confidence or clearly anomalous → isolated; doesn't even enter the review queue until it's rerun or a human steps in.
I started out using the same model to both answer and review. The agreement rate looked great. Then I realised what it was: same-source bias. A model naturally agrees with its own judgement, so a high agreement rate only proves it is self-consistent —— not that it is right.
The fix was to bring in a stronger model from a different generation for sampled calibration, and to archive every disagreement as its own document. Two independent sets of data show why this was necessary.
After rebuilding the physics ontology I ran a full review: across 637 accepted questions, the independent review agent agreed with the pipeline 85.9% of the time, and of the 90 disagreements 79 were ruled pipeline misclassifications. The important part is that those 79 were tightly clustered on 4 knowledge boundaries —— standing waves, emf and circuits, scattering and accelerators, vectors and SUVAT. Systematic errors in clusters like that never surface in spot checks; only a full cross-check brings them out.
Later, while raising quality across chemistry / biology / economics / psychology, I made stronger-model calibration a fixed step specifically to offset that bias. It caught 20 over-corrections by the review agent —— psychology's over-correction rate reached 25.8%, and one unit was ruled against 16 out of 16. Without switching models, those 20 would have been written into the bank as fixes. 164 confirmed corrections were written back in the end.
Six Edexcel IAL subjects, essentially all carrying their original figures. CIE and AQA have a classification pilot running and are labelled "in progress" inside the product.
| Subject | Questions | Status and notes |
|---|---|---|
| Mathematics | 2,649 | includes the teacher method-tag dimension for FP1/FP2/FP3 (decoupled from chapter, used only by natural-language building) |
| Chemistry | 1,056 | 0.9% residual contamination after the quality pass |
| Physics | 556 | ontology now on v3 (44 → 27 nodes) |
| Economics | 353 | essay questions need their dotted answer lines stripped before cropping |
| Biology | 317 | — |
| Psychology | 308 | — |
| Total | 5,239 | plus a dedicated multiple-choice set and a quick-drill module |
Official chapters are the exam board's view, filed against the specification. Teachers organise by "how I teach it, which method it needs". So alongside the chapter dimension I added teacher method tags as a second building dimension —— decoupled from chapter, serving natural-language building only. Because teachers often don't build a paper by chapter; they build it by method ("let's see whether they can do substitution").
The main move in physics ontology v3 was convergence: 44 nodes cut to 27. Teachers won't use a taxonomy that's too fine either —— they don't carry that many boxes in their head. After the rerun, only one unit's review rate genuinely dropped and the rest held flat, which says the extra granularity was producing noise, not precision.
What the bank stores are per-question image fragments. Turning those into a directly printable paper takes every step below, all of it hand-built.
Cut image fragments out of the source PDF question by question; questions spanning pages have to be stitched. Essay subjects like economics need their dotted answer lines stripped first, or the crop comes out as a field of dots.
Original numbers (Q3 / Q1 / Q10) are renumbered 1–4 on the new paper; multiple-choice and structured sections are numbered independently; the original number stays in the source line.
The watermark is composited into the page at render time rather than layered on top, so it can't be trivially removed.
Source footers, rough-work areas and "rough working" bands are filtered out at crop time and never reach the new paper.
Question paper and mark scheme render into two separate PDFs. The mark scheme carries official mark points (B1 / M1 / A1).
U:e3e17e38 · WS:65efa3b3 · B:5c9d8b42 · UTC —— user, paper ID, build number, generation time.
code2Session → JWTprisma migrate deploy; db push is bannedThe cross-subject leak is covered above; here are three more. The third one still isn't fixed, and I'm writing it down as it stands.
Symptom: the dedup script reported 412 duplicates.
Root cause: dedup compared full_question_text, which is a page-level field —— several sub-questions on the same page all get the same block of text. The crops were completely correct. I nearly reran the entire pipeline over this.
Rule: a field's granularity is its semantics. Before deduping, confirm whether the field describes a page or a question.
Symptom: "request failed" everywhere on a real device.
Root cause: I was testing through the admin role impersonation feature, and the impersonated identity used an ID with no real entity behind it —— so every server-side ownership check rejected it, exactly as designed.
Rule: device regression testing must use real accounts; impersonation is for UI preview only. The surface error and the real root cause frequently live on different layers.
Symptom: the same test batch is sometimes all green, sometimes red, and the victim file differs every time.
My original diagnosis: supertest spins up a temporary server per request, and port reuse creates a race. I built the "persist one global server" fix on that theory, and everything went green at the time.
Later overturned: vitest.config.ts says fileParallelism: false in plain sight —— test files don't run in parallel, so the race story doesn't hold. And that fix was never merged into the current branch. Seven independent reruns still reproduced it 3 times. It now looks more like cross-file leakage from a vi.doMock('../config/env.js') somewhere.
Rule: a plausible root cause is not a correct root cause; without falsification you haven't located anything. Also, log intermittent and deterministic failures separately —— with no local Postgres, dev-login fails every single run, and mixing that in with the flaky ones destroys your ability to judge a test run at all.
Classification quality has gates, review, calibration and internal agreement rates. But "is this the paper the teacher wanted" still rests on constraint satisfaction plus teacher review —— there is no formal eval set. Next step: a 20-paper teacher-scored set, to move "effectiveness evaluation" from intention to practice.
Different teachers will genuinely disagree on how to classify the same question. Today that goes through review_queue for a teacher to settle, but I've never computed IAA / Kappa.
The percentile definition is settled, but backfilling runs in batches. Newly loaded questions have no difficulty value until that unit's distribution is recomputed.
No CI/CD; single-point database; file storage is still local with COS / WeChat cloud as placeholders; the video path isn't deployed; the Mini Program hasn't been submitted for WeChat review.