AgentCodec (NeurIPS 2026) · Reliability for LLM and VLM agents

Make LLMs and VLMs more reliable using communication-theoretic techniques.

Animated visualization: independent prismatic channels recombining into a single reliable signal.

AgentCodec sits in front of your LLMs and VLMs and applies advanced reasoning: combining multiple models, cross-checking their answers, and adapting effort to the prompt, so answers are more accurate and consistent. You configure it once; your app just calls a single endpoint (optionally OpenAI-compatible). AgentCodec is coming soon as part of AptAgents.

How it works

Advanced reasoning, made routine.

01

LLMs & VLMs

Text or images in, text out. Vision requests route only to your vision models.

02

Combine models

Blend several models and make their reasoning deeper: ensembling, refinement, cross-checks.

03

One endpoint

Your client just makes an API call. No SDK, no orchestration code to maintain.

Inside AptAgents

Watch it run.

This is AgentCodec inside the AptAgents platform, where a reliability technique wraps your LLMs and VLMs behind one endpoint your app calls like any other model API. No SDK to adopt, no orchestration code to maintain.

Narrated walkthrough. Each endpoint puts one or more LLMs or VLMs behind a single workspace key. The Benchmarks tab scores every technique on standard or private datasets, with bootstrap confidence intervals and paired significance tests, then helps you deploy whichever one wins or train an adaptive router.
Why it matters

Reliability you can budget for.

On our paper benchmark, with one specific lineup (Nemotron and Devstral as generators, GLM-5.1 as judge), routing the technique adaptively per prompt reached ~56% cost reduction at matched quality. On quality it reached up to ~26% improvement over the single-model baseline, and ~7% at matched cost over the best fixed method we compared against for that lineup. A single knob, λ, slides the operating point along the cost/quality frontier. The qualitative pattern (adaptive beats fixed) should generalize; absolute numbers are lineup-specific.

AgentCodec cost and quality Pareto frontier benchmark
Cost and quality Pareto frontier with 95% confidence rectangles. The SemKNN router (arrows) moves the operating point toward the upper-left frontier: higher quality at lower cost than fixed prior methods.
~56%
Cost reduction at matched quality
~26%
Quality gain over a single model
28
Reliability techniques
29
Dispatchable (with baseline)
3
Adaptive routers
Vision results (New)

87.2% on MMMU, with images in the loop.

The same coding schemes run over vision-language models inside AptAgents. Swept across the MMMU multimodal benchmark with qwen3.5 and minimax-m3 as the vision-language models, the techniques trace an accuracy frontier that tops out at 87.2%, above every single-model baseline in the sweep. As with the text benchmark, the absolute numbers belong to this lineup; the pattern that governed decoding beats a single pass is what carries over.

Cost against accuracy on MMMU, one point per reliability technique, with the accuracy Pareto frontier marked
Cost against accuracy on MMMU, one point per technique, scored by exact match against the reference with no judge in the loop. The green line is the accuracy Pareto frontier, gray points are the ungoverned baselines, and the orange diamonds are cross-validated router estimates rather than measured calls.
87.2%
Peak accuracy on MMMU
2
Vision-language models combined
Exact match
Scored against the reference, no judge
Change one import

Reliability without a rewrite.

If you already have working code against the OpenAI, Anthropic, or Ollama SDK, the fastest path to reliability is one import swap. Your messages, tools, stream=True, async/sync, and the full native response shape all keep working.

Swap the import, then opt in with one kwarg
- from openai import OpenAI
+ from agentcodec.openai import OpenAI

client = OpenAI(api_key=KEY, reliability="harq_ir")   # one kwarg

resp = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "What is QUIC?"}],
)
print(resp.choices[0].message.content)        # native OpenAI shape
print(resp.reliability.technique_used)        # reliability escape hatch
print(resp.reliability.cost_usd)
The codec

6 families. 21 techniques.

21 communication-theoretic methods span these 6 families, each a coding scheme over the noisy channel that is an LLM. They generalize self-consistency, self-refine, and chain-of-verification, and outperform them on our benchmark.

MRC · SC · EGC

Diversity Combining

N parallel branches across different models (spatial), prompt variants (frequency), or temperatures (time), combined by maximal-ratio, selection, or equal-gain. Generalizes self-consistency, best-of-N, mixture-of-agents.

Chase Combining · Incremental Redundancy

Hybrid ARQ

Retry until quality clears the threshold. HARQ-IR adds new critic hints each round; HARQ-CC soft-combines every attempt into one quality-weighted answer. Generalizes self-refine and Reflexion.

Iterative SISO

Turbo Decoding

Generator drafts, critic returns structured extrinsic information, generator re-drafts on that feedback until convergence. Generalizes generator–critic and self-critique loops.

Rateless sampling

Fountain

Keep generating samples until the judge is satisfied. Easy prompts finish fast; hard prompts get more samples automatically. Generalizes universal self-consistency and dynamic-N voting.

Code rates 0.75 / 0.50 / 0.33

Forward Error Correction

A systematic block code: a primary answer plus parity sections (reasoning, verification, alternatives). The decoder cross-checks for consistency and reconstructs the corrected answer. Generalizes chain-of-verification.

Adaptive coding & modulation

ACM + SemKNN Routing

Route per-prompt by estimated difficulty. The SemKNN router learned which technique wins for which kind of prompt; a single knob λ trades quality for cost along the frontier.

Also implemented as baselines

7 prior-method baselines bring the reliability set to 28; with an uncoded single-pass baseline, that's 29 dispatchable techniques in all. Benchmark them head-to-head:

Self-Consistency Self-Refine Chain-of-Verification Best-of-N Weighted Best-of-N CISC Mixture-of-Agents
The recommended router

SemKNN: cost-aware, per-prompt.

SemKNN learned from benchmark data which technique works best for which kind of prompt. At inference it encodes your prompt locally with a small encoder model and sends only the resulting unit-norm embedding to the backend and gets back a technique recommendation.

One knob, the whole frontier

  • λ = 0Pure quality: ignore cost entirely.
  • λ = 1Balanced operating point.
  • λ = 5~10% cheaper picks at near-matched quality.
  • λ = 10~30% cheaper.
  • λ = 20~45% cheaper.

Privacy & data flow

  • Prompt text is encoded locally; only a 384-d unit-norm embedding leaves the client.
  • Sent alongside a scalar λ and a small fingerprint of your model lineup, and nothing else.
  • All other routers (fixed, acm_table, acm_linear) run fully locally, no backend at all.
  • The SemKNN backend is hosted as part of AptAgents; enterprise customers can run it inside their own network by arrangement.
Cost transparency

Every dollar is labeled.

AgentCodec never hands you a cost number without telling you how it got there. Each estimate carries a tier from a fixed enum, ranging from exact per-token rates you configured down to a parameter-count guess from the model name, so you always know how loose the accounting is.

Tiered estimates

exact_user_rate → openrouter_rate → fuzzy → table → inferred. Lower tier = tighter accounting. The worst tier across all calls is surfaced on the result.

Full per-run trace

return_trace=True gives input/output/thinking tokens, per-call cost, judge cost, number of LLM calls, and the cost-source breakdown.

Live rate catalog

Unknown models are priced against the OpenRouter catalog (disk-cached) using exact and fuzzy matching, with explicit caveats when caching or volume discounts aren't modeled.

Built for developers

A library, not a framework.

No new agent abstraction to adopt. Drop it into the SDK you already use, keep your code, and get reliability, evaluation, and cost accounting as a thin layer underneath.

One-line drop-in

Swap `from openai import OpenAI` for `from agentcodec.openai import OpenAI`. Same messages, tools, streaming, and native response shape. Add reliability with a single kwarg. Anthropic and Ollama too.

Native async streaming

27 of 29 techniques stream natively end-to-end. Per-token deltas carry role tags so you can demux drafts, critiques, and the final answer live.

Evaluation + CI gating

Compare configs with paired statistics, score against references, and wire a quality/cost gate into CI to block regressions before they ship.

Library, CLI, and FastAPI

Use it as a Python module, the `agentcodec` console script for batch JSONL runs, or behind a FastAPI / WebSocket endpoint. Construct from YAML or a dict.

Privacy by design

SemKNN encodes your prompt locally with a small encoder model and sends only the unit-norm embedding. Telemetry is anonymous and opt-out; fully on-prem deployments are available.

Provider-agnostic

Frontier APIs, OpenRouter, or local models via Ollama / vLLM / SGLang. Mix model families in one lineup: uncorrelated errors are exactly what diversity combining wants.

Under the hood

The technical picture.

AgentCodec wraps every model call in a technique drawn from classical communication theory: HARQ, diversity combining, FEC, turbo decoding, and the cost-aware SemKNN router. 28 reliability techniques (21 communication-theoretic methods across 6 families plus 7 prior-method baselines), an uncoded baseline for reference (29 dispatchable in all), and 3 adaptive routers on top. In AptAgents, it serves both language and vision-language models behind a single endpoint.

Availability

Coming soon as part of AptAgents.

AgentCodec ships as the reliability layer of AptAgents, our hosted agentic platform, which is coming soon. It works with both language and vision-language models, behind one endpoint your app calls like any other model API. The paper behind it has been accepted at NeurIPS 2026, and the preprint is available on arXiv.

What you get with AptAgents
  • Text or images in. Vision requests route only to your vision models.
  • One OpenAI-compatible endpoint behind one workspace key.
  • Several models per endpoint: ensembling, refinement, cross-checks.
  • Every technique benchmarked on your own data before you deploy the winner.

Hosted by default. Enterprise customers can run it inside their own network by arrangement.

Bring AgentCodec to your agents.

AgentCodec is coming soon as part of AptAgents: communication-theoretic reliability for your LLMs and VLMs, behind one endpoint. Tell us about your use case, and we will walk you through it.