LLM-as-a-Verifier

Open-source framework for LLM verification, agent evaluation, Best-of-N selection, and fine-grained progress tracking.

LLM-as-a-Verifier is an open-source Python framework for evaluating, verifying, ranking, and selecting outputs from large language models and AI agents. It is a framework rather than an MCP server, a single skill, a plugin, a workflow, a harness, or an AI assistant. The official primary source is https://github.com/Shubhamsaboo/llm-as-a-verifier. The README also links the project documentation at https://llm-as-a-verifier.com/docs/ and the research paper LLM-as-a-Verifier: A General-Purpose Verification Framework at https://arxiv.org/abs/2607.05391. The reviewed repository state is commit 8db8a114355a9d7fdf9a8d1d5c87f6aeebd18770, and the Python project metadata reports version 0.2.0.

What is LLM-as-a-Verifier?

The framework uses an LLM not only to generate answers but also to assess results or complete agent trajectories against defined criteria. A coding agent may produce several possible solutions for one task. LLM-as-a-Verifier can compare those candidates, calculate scores, and select a preferred solution. The input does not have to be a finished answer: observed agent steps and the state reached during a run can also be evaluated. This makes Agent Evaluation, LLM Evaluation, and verification the natural classification, while Skill Road uses the existing skill model to represent the framework because no separate framework record type exists.

Best-of-N and pairwise comparison

The public select API accepts a problem, multiple candidates, and evaluation criteria. Its result includes a best-candidate index, scores, and a ranking. Best-of-N LLM selection can be used for multiple agent runs, alternative coding solutions, or repeated rollouts of the same task. The selection is built on pairwise scoring: compare directly evaluates two candidates and returns fine-grained rewards for both. Criteria can cover correctness, root-cause analysis, or whether an agent actually verified its change. The framework does not replace the criteria: it evaluates candidates against the task and the review dimensions supplied by the team.

Fine-grained verification

The central technique goes beyond a yes-or-no label or one integer score. According to the README, LLM-as-a-Verifier reads the probability distribution over possible score tokens produced by the verifier model and calculates an expectation. The current implementation uses a 20-level letter scale from A to T for the scoring response. Multiple criteria and repeated evaluations can then be averaged. This provides a more continuous signal for choosing and analyzing candidates. It remains a model-derived signal rather than proof that an answer is factually correct, safe, or complete.

Progress tracking and online progress tracking

The track function scores a finished agent trajectory at selected checkpoints. The online ProgressTracker is designed for a still-running agent: feed it each new step, score only the prefix visible so far, and receive a progress value after each update. The verifier cannot inspect future steps. In practice, this can help teams detect poor agent runs early, abandon hopeless rollouts, resample new variants, monitor progress, and allocate compute more deliberately. The repository also demonstrates attaching images or camera frames to steps. A low score is still an input to a control decision, not an automatic reason to discard a run without review.

Test-time scaling and the Probabilistic Pivot Tournament

Test-time scaling in this project means generating multiple candidates or agent trajectories and spending additional model calls on verification and selection. More candidates, criteria, and repeated evaluations may make selection more robust, but they also increase model, token, latency, and infrastructure costs. The project implements a Probabilistic Pivot Tournament for ranking. A complete round-robin would compare every candidate pair. PPT first performs a ring pass, chooses a small pivot set, and then scores non-pivots against pivots plus pivot-versus-pivot pairs. Every candidate therefore need not be fully compared with every other candidate. The README and source describe this as O(Nk) rather than O(N²) for a fixed pivot budget; actual quality still depends on the candidate pool, pivot count, criteria, randomness, and verifier behavior.

Multimodal verification

The current public entry points accept images as a single path, URL, byte sequence, or list of those inputs. The implementation supports PNG, JPEG, GIF, and WebP detection and can load local files, HTTP URLs, or raw bytes. This enables visual evidence such as robot rollout frames, before-and-after screenshots, UI states, or camera images to participate in agent trajectory evaluation. Multimodal verification remains dependent on the selected model's vision capabilities and on the relevance and quality of the supplied images.

Supported models and backends

The package depends on google-genai and openai. For Gemini, the current implementation documents Vertex AI with VERTEX_API_KEY because extracting logprobs requires the Vertex API path. DeepSeek is available through its hosted API with DEEPSEEK_API_KEY, and the current README names deepseek-v4-flash as the verifier in the self-verification benchmark. The framework also accepts OpenAI-compatible servers that return token-level logprobs. The README shows vLLM serving Qwen/Qwen3.5-9B with OPENAI_BASE_URL as an example. The source also refers to SGLang, OpenAI, and DeepSeek in the compatible client path. An optional vLLM extra requires vLLM 0.19 or newer. A benchmark model is not automatically a fully supported verifier integration: GPT-5.5, Opus, Gemini 3 Flash, and Claude Opus 4.8 appear as base models in benchmark configurations, while the verifier backend has separate logprob requirements.

Coding agents and benchmarks

The framework is particularly suited to coding agent evaluation and agent trajectory evaluation. The current README documents Terminal-Bench 2.1, SWE-Bench Verified, and MedAgentBench with included trajectories and reproduction scripts. Its expected-results table reports 86.5 percent for Terminal-Bench V2, 78.2 percent for SWE-Bench Verified, and 73.3 percent for MedAgentBench for LLM-as-a-Verifier. According to the README, Gemini 2.5 Flash is used as the verifier configuration across the listed benchmarks, while base models, harnesses, and Best-of-N settings differ by benchmark. These are project-reported results from the reviewed repository state, not independently confirmed performance guarantees from Skill Road. The README also names RoboRewardBench as an additional application area without providing the same result table in the reviewed repository.

Relationship with TurboAgent

TurboAgent is a separate repository at https://github.com/llm-as-a-verifier/TurboAgent. Its README describes a Claude Code plugin and LLM API proxy that generates multiple responses concurrently and uses LLM-as-a-Verifier with the Probabilistic Pivot Tournament to select one. TurboAgent is not part of this catalog record, and its proxy endpoints, provider configuration, and visualizer are not presented as properties of the framework itself. Skill Road does not create a separate TurboAgent record in this publication because the request concerns LLM-as-a-Verifier and the current skill schema has no project-relation table.

Audience, requirements, and costs

The framework fits AI agent developers, autonomous coding-agent developers, AI engineers, researchers, evaluation-system and multi-agent developers, and teams operating their own agent harnesses. Requirements include Python 3.9 or newer, pip installation, and a verifier backend with the required logprob capability. Hosted providers need their relevant API key; a local OpenAI-compatible server is an alternative. The framework itself is open source under the MIT License and available without a Skill Road product charge. External model and token charges can still apply, and local deployments incur their own compute and infrastructure costs.

Benefits, limitations, and license

Documented benefits include automated evaluation of agent outputs, Best-of-N selection, fine-grained scores, online progress monitoring, multimodal input, Python integration, and multiple backend paths. Important limitations remain: verifier models can be wrong, additional candidates increase cost, criteria shape the result, and a high verifier score does not guarantee factual or secure correctness. LLM evaluation does not replace unit tests, integration tests, static analysis, CI, security review, or human approval. The LICENSE file identifies the project as MIT licensed. This is a factual license summary, not legal advice.

Catalog classification

Skill Road classifies this project as a framework in the main Development category. Agent Evaluation, LLM Evaluation, and Verification are the editorial subtopics. Coding and Coding Agents remain useful tags and application areas, but they are not the main category because the framework also evaluates general agent trajectories and multimodal scenarios.

Conclusion

LLM-as-a-Verifier is an agent evaluation framework for LLM verification, LLM output ranking, coding agent evaluation, agent trajectory evaluation, Best-of-N LLM selection, test-time scaling, and AI agent verification. It combines pairwise comparison, probabilistic score expectations, pivot-based ranking, online agent progress tracking, and multimodal inputs. Its strongest role is to add structured evidence for selection and monitoring; it should not be treated as a guarantee of correct agent results or as a replacement for deterministic engineering tests.

Free
Provider
Shubhamsaboo
License
MIT
Last reviewed
10.09.2026

Repository and documentation

Categories