
We tested five small, low-cost models on Litewave AI’s batch-review benchmark of real-world compliance checks against a real-world executed batch record plus an expert answer for each check. Qwen3.8 Flash has the best accuracy of 79% of checks answered correctly according to the expert, with evidence. DeepSeek V4 Flash (another open-source model) places second at 74% accuracy. Evidence, more than just correctness of verdicts, is what separates the models.
Each model is evaluated on a set of already digitized batch records, not a scanned image. Each record is represented as text, structured in a way that a computer can parse, and the set includes all the batch records produced by a manufacturing site: manufacturing records, records specific to quality control and quality assurance, and ERP data. The challenge of converting a scanned document into this structured digital format is an important task but is outside the scope of this benchmark.

Why batch review is hard
Each batch of an active pharmaceutical ingredient ends in an executed batch record: scores of pages of handwritten times, temperatures, weights, signatures, and laboratory results. Before the batch can be released, a reviewer reads it against the protocol the quality organization wrote for that batch. Every temperature within range for the required time, every signature, every analytical result within spec, and every yield reconciled.
This is not a paperwork exercise. Of the 2,837 citations FDA wrote to drug manufacturers in fiscal 2025, the regulation covering production record review and the investigation of discrepancies, 21 CFR 211.192, was the second most cited: 236 of them, one finding in every twelve.¹
A check is not only about reading the right rows, it is about finding all of them. A rule that covers every operation in a range, or every entry in a table, fails the moment one in-scope row is never considered: a model that reads nine rows of ten perfectly has still missed the finding.
A language model can read the record. The question for an audit tool is whether it can demonstrate its work. A judgment without the reading the page and the limit is not a guarantee, and a reviewer cannot be expected to sign off on conjecture.
Three failure patterns emerged as soon as models were tested on the task:
- •Right verdict, wrong evidence. The model reports a discrepancy but cites the wrong row, or gives the wrong reason.
- •Invented readings, data, and processes. A value that is not in the record, a page that does not hold it, a legible row called unreadable.
- •Variance in outputs across runs. The same rule and the same record give a different outcome on a rerun.
This benchmark measures all three, across 5 different models.
How the benchmark works
The dataset
Real rules and real records. Each check is a pair between a rule written by a pharmaceutical quality team and an executed batch record from a site manufacturing an active pharmaceutical ingredient. The model is fed the record as a text extracted from the scanned pages and the rule as a short statement of what to be checked, on which document and against which limit. A single check is handed between about 15 and 183 pages of record, a median of 45, and around 12% of checks read more than one document. The full record sets behind them are larger still, up to around 250 pages.
Discrepancies and clean records. About 29% of checks have a real discrepancy and the rest are clean, as in routine reviews. A model that always answered “no discrepancy” would get about 71% of verdicts correct, and score zero here, since it never provides evidence.

Scoring
Each check has an expert-curated answer: the discrepancy verdict (present or not present), and the evidence, cited data, performed checks as additional support. The expert answers were developed with the help of pharmaceutical quality assurance and quality control experts.
Every expert answer was reviewed against its record. Roughly a quarter of them were corrected based on the source data, six in the verdict itself. This also shows that even human experts can make mistakes when performing batch review.
Note: all scores here are based on the corrected answers, and the principles they follow are in the appendix.
Two things are scored on every answer.
- 1.The verdict. Does the model’s “discrepancy present” or “none detected” match the expert’s? A refusal or “cannot verify” counts as wrong.
- 2.The evidence. An AI grader compares the model’s answer with the expert’s, check by check, against a fixed rubric, and gives an evidence score from 0 to 10. A score of 8 or more is counted as a pass.
Note: the appendix walks through one check scored end to end.
Overall accuracy is computed across both axes on the same answer: the right verdict and passing evidence, a grader score of 8 or more. This allows us to test both the classification capabilities of the model and its ability to gather and present evidence, both very important qualities in the field of regulation and compliance.
We also report other secondary metrics:
- •Precision and recall on real discrepancies, where a detection counts only if its evidence passes.
- •Fabrication rate, the share of answers that state something the record does not support.
- •Consistency across 3 independent runs of the full benchmark.
Results
Qwen3.8 Flash leads, based on both discrepancy classification, and evidence

Qwen3.8 Flash in first reaches 79.4% overall accuracy, 12 points ahead of GPT-5.6 Luna, which falls in last place. Requiring the evidence to hold, and not only the verdict, costs every model 10 to 18 points, and shifts the order in the middle of the table. DeepSeek V4 Flash and Gemini 3.5 Flash Lite finish within 1 point of each other, tighter than this benchmark can separate.
Precision and recall consider only real discrepancies. Qwen3.8 Flash bests them reliably, with a precision of 69%, while Gemini 3.5 Flash Lite has the lowest recall, 44%, catching fewer than half of real discrepancies with sound evidence.
| model | overall accuracy, range over runs | evidence score, 0 to 10 |
|---|---|---|
| Qwen3.8 Flash | 79.4% (78.7% to 79.8%) | 8.8 |
| DeepSeek V4 Flash | 73.7% (73.3% to 74.2%) | 8.3 |
| Gemini 3.5 Flash Lite | 72.8% (70.9% to 73.8%) | 8.3 |
| Gemma 4 26B A4B | 67.9% (66.1% to 70.4%) | 8.1 |
| GPT-5.6 Luna | 67.5% (67.0% to 67.7%) | 7.9 |
Where every model struggles

The hardest families of batch review checks for every model ask for multiple readings across an entire record, or for values reconciled across tables: cross-document 47%, batch record review 49%, in-process results 54%, averaged over the models. One incorrect or missing reading is grounds for failure in the evidence check. Checks that read two to four documents at a time cost every model 15 to 38 points of accuracy over single-document checks.
Length is the other lens to evaluate these models. Every model in the table scores lower on checks over 64,000 tokens than on checks under 32,000, by 5 to 16 points, and Gemma 4 26B A4B falls the furthest. Long records are the normal case in the batch review, not the exception. Exact numbers are in the appendix.
Consistency and explainability
A reviewer needs two things besides just a correct verdict: consistency across multiple executions of the same batch. If a batch is compliant today, under the same parameters, it should also remain compliant tomorrow.
Each model ran the full benchmark 3 times with the same prompt and the same records.
| model | passes in every one of 3 runs | outcome changes between runs | evidence holds when there is a discrepancy to explain |
|---|---|---|---|
| Qwen3.8 Flash | 70% | 17% | 78% |
| DeepSeek V4 Flash | 58% | 29% | 78% |
| Gemini 3.5 Flash Lite | 63% | 18% | 63% |
| Gemma 4 26B A4B | 56% | 24% | 62% |
| GPT-5.6 Luna | 57% | 20% | 70% |
Qwen3.8 Flash passes 70% of checks every time and Gemma 4 26B A4B passes 56%, while between 17% and 29% of checks change outcome from one run to the next.
The last column is the explanation. When the verdict is already right, the evidence behind it holds about nine times in ten on a clean record, but only 62% to 78% of the time when there is a real discrepancy to describe. Saying that something is wrong is the easy half. Saying which row, which reading and which limit is the half a reviewer acts on.
Across every model and all 3 runs, only 4% of checks are never passed. Almost every check in this benchmark is within reach of a small model on some run; what is missing is getting it right the same way every time.
Fabrication is a model trait
An independent blind review looked at one answer for each check against the whole record, and counted the answers that state something that the record does not support.
| model | fabrication rate |
|---|---|
| GPT-5.6 Luna | 12% |
| Qwen3.8 Flash | 15% |
| Gemini 3.5 Flash Lite | 23% |
| DeepSeek V4 Flash | 26% |
| Gemma 4 26B A4B | 30% |
The gap between GPT-5.6 Luna and Gemma 4 26B A4B is bigger than their gap in overall accuracy. For an audit tool a model that invents in one out of every three answers is a fundamentally different product than one that does it in one out of every eight.
What we take from this
None of these models are capable of reviewing a batch on its own
Even the best among them fails about one check in five. In a regulated release workflow, that is a check a reviewer has to redo.
About half the failures are evidence failures, not judgement failures.
In 49% of Qwen3.8 Flash’s failed checks the verdict was right and the supporting reading, row or limit was wrong or missing. The model reached the correct conclusion and could not show why. That points at how the record is given to the model and how its answer is checked, not at model size.
Open weights lead. Qwen3.8 Flash tops the benchmark, and DeepSeek V4 Flash is level with the best closed model. Open-weight models can run inside a manufacturer’s own environment, so batch records never leave it.
Appendix
A worked example
One check from the benchmark, with identifiers removed. It shows the rubric and why the correct verdict is not sufficient.
Rule/Batch Review Check given to model
Verify that the parameter record sheet records at least 4 temperature and vacuum readings for operation 93, spaced 20 to 40 minutes apart, with temperatures below 50 °C and vacuums within 600 to 720 mmHg for each reading.
What the record shows on the parameter record sheet for operation 93. The first reading’s vacuum field has a handwritten note stating as written, rather than a number.
| time | temperature | vacuum |
|---|---|---|
| 11:25 | 36.1 °C | Vaccum Apply |
| 11:52 | 40.3 °C | 650 mmHg |
| 12:20 | 43.8 °C | 650 mmHg |
| 12:58 | 48.6 °C | 650 mmHg |
The expert answer, which is the rubric for this check.
Verdict: DISCREPANCY PRESENT
| check | criterion | outcome | values the expert cites |
|---|---|---|---|
| temperature | every reading below 50 °C | pass | 36.1, 40.3, 43.8, 48.6 °C |
| vacuum | every reading within 600 to 720 mmHg | fail | the first reading has no numeric value; the other three read 650 mmHg |
| interval | consecutive readings 20 to 40 minutes apart | pass | 27, 28 and 38 minutes |
| reading count | at least 4 readings | pass | 4 timed readings |
Answer A: right verdict, sound evidence. Evidence score 10, passes. It reports a discrepancy because the first reading’s vacuum was not a number but rather a phrase, “Vaccum Apply”, hence the reason why vacuum compliance cannot be shown for all four readings. It cites the three vacuum readings in numbers, all four temperatures, and the 27, 28, and 38 minutes. The grader finds all the expert check in the answer, and all the numbers match.
Answer B: right verdict, unsound evidence. Evidence score 6, fails. It also provides a discrepancy, but blames it on the number of readings; it counts three, when there should be four. The same answer, however, checks all the vacuum readings as being within range. The grader records the vacuum check as contradicted, because the answer marks it compliant where the expert fails it, and the number of readings as contradicted, 3 against 4. The verdict is correct, but the reasoning is faulty. Someone grading using answer B would accept the vacuum readings, which is the single area where the record is wrong. Thus, in overall accuracy, Answer B is incorrect.
The judge
The evidence score is determined by an AI grader operating under a fixed rubric, rather than from a free-form rating. The grader uses DeepSeek V4 Pro at temperature 0, the same across all models, and is blind to which model answered the question. It performs two steps:
- 1.Extraction, in which it reads the expert answer and the model answer side by side, and records for every check the expert performed the expected value and outcome, the model’s value and outcome for the same check, and whether they agree, contradict, or the check was missing, and notes which check the expert’s verdict rests on.
- 2.Scoring by code: 10 points if all checks are present and agree, 6 points if a supporting check is missing or disagree, 3 points if the check the verdict rests on disagrees, and 2 points if the verdict itself disagree. Eight or more passes. The grader itself never picks the number of passes, and the rubric uses the same internal scoring.
We validated the grader against an independent blind review. A second model, Claude Sonnet 5, read every answer in this report against its record and the same rubric to determine whether it would pass. The two agree on pass or fail for 88% of answers, with Cohen’s kappa of 0.62 to 0.74 across the models.
Note: Cohen’s kappa is a statistic that corrects agreement for chance, and values over 0.6 are conventionally read as substantial.
The agreement is identical for all models and where there is a discrepancy between the two the grader is the more stringent. The grader shares a model family with DeepSeek V4 Flash; that model sits mid-field and its agreement with the independent review matches the others, hence no indication of favoritism.
How the expert answers are written
Every expert answer follows the same principles and has been checked for them; the scores given in this report are based on corrected answers.
- •A value is compared exactly as written, never reinterpreted. The model operates strictly on the data provided to it.
- •Units are part of the value. Two values match if both their number and unit match.
- •Blank or unreadable rows fail. A row within the scope of the rule that is blank, has a dash, or has text where number is expected fails the check, unless the rule says to skip such rows.
- •Every row within the rule’s scope is evaluated, and nothing outside of it.
- •Each check is a comparison of some recorded values against a criterion. Notes about method or scope are not checks and are never scored.
- •Every cited value is in the record. Every number, page and row the expert cites exists in the record exactly as cited.
- •The verdict follows from the checks.
- ◦Discrepancy Present if at least one evaluated row fails.
- ◦None Detected if every evaluated row passes.
- ◦Cannot Verify only if the record contains no row in the scope.
Overall accuracy
The verdict matches the expert’s and the evidence score is 8 or more on the same answer. A refused, unanswered or “cannot verify” answer is counted wrong, and every check in the benchmark counts, so nothing is dropped for being hard. Reported as the mean over the 3 runs, with the lowest and highest run in brackets.
Precision and recall
A detection occurs only when the model flags a true discrepancy and its evidence passes. Thus, precision is the proportion of a model’s signaled discrepancies that are true, with passing evidence, and recall is the proportion of true discrepancies caught with passing evidence. Finally, F1 is their harmonic mean, and hence is high only when both are high.
Fabrication rate
A single sampled answer was evaluated, blind, against the full record, by an independent reviewer model, Claude Sonnet 5, not used on the original set. An answer was deemed wrong if it gave a different value, quote, outcome, or page than the record or declared a legible row as illegible. Omissions, and disagreement with the expert’s verdict, were not counted.
Every model in the comparison was evaluated on the same set of checks, with the responses for a given check being presented together, but shuffled, so as a reviewer they couldn’t know which answer came from which model. This is important context, as the stringency of the review process affects the absolute rates, as we have seen the same answers reviewed at twice the rate on a separate set of evaluations. The numbers in this section should be read as comparisons between the models, not an absolute value.
Consistency
Every model ran the full benchmark 3 times with the same prompt and records. A check passes in every run, fails in every run, or is mixed.
Full results
All metrics, mean over 3 runs.
| model | overall accuracy | precision | recall | F1 | evidence score | fabrication rate | no verdict given |
|---|---|---|---|---|---|---|---|
| Qwen3.8 Flash | 79.4% (78.7% to 79.8%) | 0.69 | 0.68 | 0.68 | 8.8 | 15% | 3.4% |
| DeepSeek V4 Flash | 73.7% (73.3% to 74.2%) | 0.60 | 0.67 | 0.64 | 8.3 | 26% | 4.7% |
| Gemini 3.5 Flash Lite | 72.8% (70.9% to 73.8%) | 0.54 | 0.44 | 0.49 | 8.3 | 23% | 1.3% |
| Gemma 4 26B A4B | 67.9% (66.1% to 70.4%) | 0.45 | 0.58 | 0.51 | 8.1 | 30% | 1.6% |
| GPT-5.6 Luna | 67.5% (67.0% to 67.7%) | 0.55 | 0.63 | 0.58 | 7.9 | 12% | 12.0% |
Consistency over 3 runs. A check passes in every run, fails in every run, or is mixed.

| model | passes in every run | mixed | fails in every run |
|---|---|---|---|
| Qwen3.8 Flash | 69.7% | 17.5% | 12.8% |
| DeepSeek V4 Flash | 57.8% | 28.9% | 13.2% |
| Gemini 3.5 Flash Lite | 63.2% | 17.9% | 18.8% |
| Gemma 4 26B A4B | 56.1% | 24.2% | 19.7% |
| GPT-5.6 Luna | 57.2% | 20.4% | 22.4% |
Checks with a discrepancy and clean checks. Each accuracy is scored separately on its own set of checks. On the right of the chart, when the verdict is already right, how often the evidence holds.

| model | accuracy, checks with a discrepancy | accuracy, clean checks | evidence holds when the verdict is right, discrepancy | evidence holds when the verdict is right, clean |
|---|---|---|---|---|
| Qwen3.8 Flash | 67.7% | 85.8% | 77.7% | 93.1% |
| DeepSeek V4 Flash | 67.4% | 76.8% | 77.9% | 91.4% |
| Gemini 3.5 Flash Lite | 44.2% | 85.0% | 63.2% | 90.4% |
| Gemma 4 26B A4B | 57.6% | 72.4% | 62.4% | 86.0% |
| GPT-5.6 Luna | 62.7% | 70.2% | 69.5% | 92.6% |
By rule family, with each family’s share of the benchmark.
| rule family (share of benchmark) | Qwen3.8 Flash | DeepSeek V4 Flash | Gemini 3.5 Flash Lite | Gemma 4 26B A4B | GPT-5.6 Luna |
|---|---|---|---|---|---|
| Process record sheets (35%) | 86.8% | 79.7% | 80.6% | 82.5% | 82.1% |
| Batch record review (13%) | 68.5% | 60.1% | 46.4% | 28.6% | 42.9% |
| Analytical test reports (8%) | 79.3% | 73.9% | 89.2% | 72.1% | 60.4% |
| QC results (4%) | 73.7% | 63.2% | 68.4% | 61.4% | 33.3% |
| In-process results (3%) | 54.8% | 47.6% | 78.6% | 47.6% | 40.5% |
| Cross-document (2%) | 63.6% | 48.5% | 54.5% | 33.3% | 33.3% |
| Other batch-record checks (34%) | 80.0% | 78.0% | 71.9% | 71.7% | 72.5% |
By number of documents a check reads.
| checks over | Qwen3.8 Flash | DeepSeek V4 Flash | Gemini 3.5 Flash Lite | Gemma 4 26B A4B | GPT-5.6 Luna |
|---|---|---|---|---|---|
| single document | 79.9% | 75.3% | 74.3% | 69.9% | 70.1% |
| 2 to 4 documents | 65.3% | 56.9% | 50.0% | 45.8% | 31.9% |
| a full set of related records | 83.3% | 65.5% | 71.4% | 59.5% | 60.7% |
By prompt size, counted with one fixed tokenizer across models.
| checks by prompt size | Qwen3.8 Flash | DeepSeek V4 Flash | Gemini 3.5 Flash Lite | Gemma 4 26B A4B | GPT-5.6 Luna |
|---|---|---|---|---|---|
| under 16k tokens | 78.8% | 79.2% | 73.8% | 69.2% | 74.6% |
| 16k to 32k | 82.8% | 79.3% | 80.1% | 76.0% | 71.5% |
| 32k to 64k | 78.3% | 71.2% | 69.4% | 65.6% | 65.2% |
| over 64k | 76.1% | 63.1% | 66.2% | 57.2% | 57.7% |
Models
| model | weights | endpoint | model version string |
|---|---|---|---|
| Qwen3.8 Flash | open weights | Alibaba Model Studio | qwen3.8-flash |
| Gemini 3.5 Flash Lite | closed | Google Vertex AI | gemini-3.5-flash-lite |
| DeepSeek V4 Flash | open weights | Alibaba Model Studio | deepseek-v4-flash-0731 |
| GPT-5.6 Luna | closed | Azure OpenAI | gpt-5.6-luna |
| Gemma 4 26B A4B | open weights | Google Vertex AI | google/gemma-4-26b-a4b-it-maas |
Provider default settings and the same prompt for every model. Model versions and endpoints as of 2026-09-06 to 2026-09-15.
Limitations
The benchmark separates models differing by several points, but cannot distinguish among models differing by about two. The grader and fabrication review are AI models validated against an independent blind review, not human labels. Every model was accessed via a hosted endpoint, not self-hosted, two of which were external to our own cloud accounts; one of those two rejected a small fraction of requests after content inspection, and those count as errors. Results hold for these records, these rules and these model versions.
Disclosures
- •Litewave AI is an entity that creates software for the purpose of batch record review. None of the models evaluated below are developed by Litewave AI.
- •The grader, DeepSeek V4 Pro, belongs to the same model family as one of the models being tested, DeepSeek V4 Flash. The specific test done to assure no favoritism is described in The judge.
- •The records and rubrics are confidential information belonging to the customer and cannot be disclosed. The rubrics and the scores are posted here for reference in their entirety. The customer’s names were redacted from the prompts sent out to the models hosted outside of our cloud accounts.
Notes
¹ FDA, Inspection Observations, FY 2025 Excel file (drugs program area), accessed 16 September 2026. fda.gov/inspection-observations
