Google's research team submitted a paper on arXiv on October 1, disclosing an agentic verification framework called VeriHarness, with the code released on GitHub under the Apache 2.0 license. The problem it targets is concrete: when an agent runs a long task that takes dozens of steps and turns in a spreadsheet, a report, or an edited file, who decides whether that output is actually correct.
VeriHarness's approach is to have the same model that generated the answer act as its own reviewer. The model runs independently on the same task multiple times, producing several results, which are then laid out side by side and sorted into two buckets — "disagreement" and "consensus" — each handled differently.
Two review tracks
The first track handles disagreement. When the results disagree on a conclusion, the reviewer goes back into the task environment to look for evidence — reopening the original file, checking a specific cell in a spreadsheet — and uses whatever it can verify in the environment to rule out the wrong claims.
The second track handles consensus. Even when the results agree, the framework doesn't just wave them through. It actively looks for holes, searching specifically for evidence that could overturn the shared conclusion, while also checking whether every run missed the same requirement in the task. The paper's premise is that multiple runs landing on the same answer doesn't guarantee that answer is correct — the model could easily make the identical mistake every time.
Once both tracks finish, the process moves into a verdict stage: the reviewer picks the strongest result as a base, revises it based on the evidence it has gathered, redoes it entirely if needed, and separately logs any questions it still couldn't resolve. The framework equips the reviewer with a workspace, evidence-gathering tools, and a set of reusable "verification skills" that the paper says can improve themselves based on failure feedback.
What was tested, and how much it helped
The evaluation covers five long-horizon workspace benchmarks: APEX-Agents, Workspace-Bench Lite, WorkBuddy Bench, SpreadsheetBench 2, and JobBench — most of them tasks that require delivering a finished file, such as office documents, spreadsheets, or job-application materials. The same model handles both generation and review; two models were tested, Gemini 3.5 Flash and Claude Opus 4.8, both accessed through Vertex AI.
According to the paper and its README, VeriHarness ranks first on "selection score" across all five benchmarks, against baselines that include single-run generation and earlier LLM-as-a-Verifier methods. With evidence-driven revision added, Gemini 3.5 Flash gained an average of 6.2 points over single-run generation, and Claude Opus 4.8 gained an average of 6.4 points.
The team also released roughly 26,000 run trajectories on Hugging Face — both models run 10 times per question across the five benchmarks, complete with rendered execution traces, delivered files, and scores. The benchmarks' inputs and reference answers were not released. The paper notes that producing this dataset cost more than $100,000.
How it differs from "run it a few times and take the majority"
Running a model several times and voting on the result is a common way to squeeze out extra accuracy in the industry — cheap, easy to implement, and effective on short Q&A tasks. On long tasks it runs into two problems: the deliverable is a file, which can't simply be voted on, and when several results agree, voting preserves any shared mistake intact. VeriHarness's two tracks are aimed squarely at those two problems.
Another common approach is to have a model act as a judge, reading through the results and scoring them directly. That approach depends on the judge's reading and judgment and never loops back to check against the environment. VeriHarness puts "can the claim be backed up with evidence found in the environment" at the center, at the cost of turning the review step itself into an agentic task that calls tools and spends tokens. The paper doesn't give a figure for how much extra overhead review adds relative to generation — for now, companies that want to work out that cost have to measure it themselves.
This outlet previously covered Microsoft's ThinkingBox benchmark, which repeats the same task 20 times to see how stable an agent's results are. The two pieces of work point in the same direction: on long tasks, a single run's score has limited value, and the variation across multiple runs is itself a signal — one side uses it to measure stability, the other to catch errors.
Is it ready to use
The README specifies a Linux environment, requiring Python 3.10 or later, Node.js 22.19 or later, Docker, and unprivileged user namespaces; underneath, it runs on the open-source pi runtime, so it can connect to any model provider that pi supports. The project page also includes this line:
“This is not an officially supported Google product.”
In other words, this is research code for now, and Google isn't committing to maintain it. Neither model tested in the paper is each company's latest generation, and the authors didn't test how much of the improvement would remain with Gemini 4 or Opus 5.5 — that part will have to wait for third-party reproduction.
Sources: the arXiv paper “VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks,” the Google Research GitHub project page, CocoLoop, and the Hugging Face dataset page; the paper and README were checked for the five benchmark names, the two models' average score gains, the roughly 26,000 trajectories, and the runtime requirements.