Accuracy

How do you know it’s right?

Two things have to be true for compliance work to be useful: the model has to be smart enough to read regulations correctly, and the workflow has to make every answer checkable. This page covers both — with the actual numbers from independent benchmarks and a plain-English walkthrough of how ExChek keeps Claude honest.

First — the substrate

Claude Opus 4.7 leads on the benchmarks closest to compliance work.

Three benchmarks below test exactly the kind of work an export-compliance agent has to do every day: real knowledge work (drafting, summarizing, decisions), reading long technical documents, and reasoning across hundreds of pages of context. We didn’t pick these — Anthropic published them.

Knowledge work benchmark (GDPVal-AA): Opus 4.7 1753 Elo, GPT-5.4 1674, Opus 4.6 1619, Gemini 3.1 Pro 1314.
Knowledge work (GDPVal-AA). Real-world deliverables: drafting, analysis, decisions. Opus 4.7 leads at 1753 Elo — ahead of GPT-5.4 and Gemini 3.1 Pro.
Document reasoning benchmark (OfficeQA Pro): Opus 4.7 80.6%, Opus 4.6 57.1%, GPT-5.4 51.1%, Gemini 3.1 Pro 42.9%.
Document reasoning (OfficeQA Pro). Reading and answering questions from real business documents. Opus 4.7 hits 80.6% — nearly 30 points clear of the next model.
Long-context reasoning benchmark (GraphWalks at 1M tokens): Opus 4.7 75.1% Parents and 58.6% BFS, Opus 4.6 71.1% and 41.2%.
Long-context reasoning (GraphWalks, 1M tokens). Holding hundreds of pages of context and reasoning across all of it — the heart of compliance work. Opus 4.7 leads on both splits.

Why this matters for export compliance

These three skills are the job.

01

Knowledge work = drafting your memo.

Every classification ends with a one-page audit-ready memo — with reasoning, citations, and a named reviewer. That’s knowledge work. Opus 4.7 leads.

02

Document reasoning = reading your part spec.

A datasheet is a business document. Pulling the encryption strength, the country of origin, the technical parameters — that’s OfficeQA-style work. Opus 4.7 wins by 23+ points.

03

Long-context = reading the regulations end-to-end.

15 CFR Part 774 alone is hundreds of pages. The Commerce Country Chart adds more. ExChek loads all of it and Claude has to reason across the whole stack. Long-context is the foundation.

The full picture

Across every category, head-to-head.

Anthropic published the full comparison — agentic coding, tool use, computer use, financial analysis, multilingual Q&A, and more. ExChek runs on the highlighted column.

Full Anthropic benchmark table comparing Opus 4.7, Opus 4.6, GPT-5.4, Gemini 3.1 Pro, and Mythos Preview across agentic coding, tool use, computer use, search, financial analysis, cybersecurity, graduate reasoning, visual reasoning, and multilingual Q&A.
Source: Anthropic — Claude Opus 4.7 announcement

ExChek’s accuracy layer

A smart model isn’t enough. We make every answer checkable.

Even the best frontier model can be wrong. ExChek wraps Claude with four layers that force every output to be defensible.

Layer 1

Live eCFR data, every time.

We don’t rely on the model’s training cutoff. ExChek pulls the current text of 15 CFR 774, 22 CFR 121, and the Commerce Country Chart from the official eCFR API on every classification.

Layer 2

Citation-backed determinations.

Every memo includes the exact CFR section, paragraph, and date. If we say a part is EAR99, we show you the rule and why nothing else applied.

Layer 3

Human-in-the-loop, always.

ExChek never auto-files anything. You read the determination, the reasoning, and the citations — and you sign off. Your name goes on the memo, not Claude’s.

Layer 4

HMAC-chained audit log.

Every decision is hashed and chained. If anything is altered after the fact, the chain breaks. Tamper-evident by default — the kind of trail BIS expects on a five-year recordkeeping requirement.

Try it yourself.

Download the plugin, run it on a part you actually ship, and read the memo. Then check it against the CFR yourself. That’s the only accuracy test that matters.