A Full-Paper Reading Test: Correct Tables, Two Interpretation Errors
October 2, 2026 3 min read
A one-off full-paper reading exercise in SentX ended with a mixed result: the response reproduced the requested table figures accurately, but it made two substantive interpretation errors. All 25 pages of the public CC BY 4.0 MIST preprint — "Recalling Too Well: Sycophancy Evaluation and Mitigation in Memory-Augmented Models," arXiv 2606.10949v2 — were extracted to plain text and supplied in a single message. The prompt named the tables and appendices to inspect; separately, reviewers had frozen nine reference topics and their expected answers before the request, and those checks were not shared with the model. One original reply through normal SentX defaults was retained as given, with no selected rerun after errors appeared. The paper evaluates its own named memory systems and response models; it did not study SentX or Victoria.
Because the input was text only — not a PDF upload, image, or figure — the exercise covers reading of supplied full text, not visual or figure reading, file parsing, autonomous discovery, or cross-session recall.
What the response got right was concrete. It reproduced the requested Table 3 comparison exactly: GPT-5.2 on MIST-Moral at 5.7±0.8% strict sycophancy with chat history versus 41.2±1.5% with Mem0 across five runs — a gap of +35.5 percentage points, approximately 7.23 times, or +623% in relative terms. It reproduced all eight reported means in Table 6, from a separate three-run experiment, with their displayed ± values. It distinguished strict sycophancy from overall error rate, and it identified the Appendix L per-question comparison that came out non-significant under adjustment (p=0.141) despite a run-level p=0.033.
Two substantive interpretation errors survived. First, the response claimed increases for all memory-system/model pairs. Table 7 shows why that overreaches: MiniMax 2.5 on MIST-Science moves from 5.7±1.0% with chat history to 4.9±0.6% with MemOS. The summary generalized from a dominant pattern to a universal claim.
Second, it read the abstract's "up to 40%" as a precise relative-percentage claim. The abstract does not define a single named comparison establishing that reading — percentage points and relative percent are different units, and the wording stays ambiguous in the source. The defensible move is to retain the ambiguity rather than supply an interpretation the text never establishes.
Smaller notes round out the review: the realism aggregation treatment, the fact that the Table 6 caption does not explicitly label the ± values as standard deviations, and an omitted metric-denominator detail. None changes the headline findings, but each is the kind of item that reads fine until someone checks the row.
This is one first-party demonstration, not a benchmark. Nothing in it establishes a general accuracy score, an independent evaluation, or product superiority — and successful numeric reproduction does not establish understanding. Both errors were interpretive, sitting exactly where the numbers looked right.
For researchers, the useful takeaway is a short checking routine rather than a disclaimer. Ask the assistant about named tables and metric denominators, not just results. Compare every summary claim against the original rows, deliberately hunting exceptions and counterexamples. Retain the first answer verbatim so corrections can be diffed against it, and save a corrected, source-linked note alongside. That is also how SentX's public research summarizer workflow positions itself: paste text or upload a PDF, query methods, results, and limitations, and check important claims against the source. In this case, the same pass that confirmed Table 3 would have caught both errors — one by surfacing the contradicting row in Table 7, the other by refusing to resolve the abstract's units.