Make LLMs and VLMs more reliable using communication-theoretic techniques.
AgentCodec sits in front of your LLMs and VLMs and applies advanced reasoning: combining multiple models, cross-checking their answers, and adapting effort to the prompt, so answers are more accurate and consistent. You configure it once; your app just calls a single endpoint (optionally OpenAI-compatible).
Advanced reasoning, made routine.
LLMs & VLMs
Text or images in, text out. Vision requests route only to your vision models.
Combine models
Blend several models and make their reasoning deeper: ensembling, refinement, cross-checks.
One endpoint
Your client just makes an API call. No SDK, no orchestration code to maintain.
Watch it run.
This is AgentCodec inside the AptAgents platform: the advanced build, where a reliability technique wraps your LLMs and VLMs behind one endpoint your app calls like any other model API. No SDK to adopt, no orchestration code to maintain.
One codec, two builds.
The research library and the platform build share the same coding schemes. They differ in what they can see, and in how much you have to wire up yourself.
Source-available library
The published research build. Install it from source, point it at your own provider keys, and run the coding schemes in your own process.
- Text in, text out. Language models only.
- 28 reliability techniques and 3 adaptive routers, all configurable from YAML.
- One-line drop-in for the OpenAI, Anthropic, and Ollama SDKs.
- Free for research, teaching, and internal noncommercial evaluation.
Platform build
The advanced build, and the only one that sees images. Configure an endpoint once in the console; your app then calls it like any other model API.
- Text or images in. Vision requests route only to your vision models.
- One OpenAI-compatible endpoint behind one workspace key. No SDK to adopt.
- Combine several models per endpoint: ensembling, refinement, cross-checks.
- Benchmark every technique on your own data, then deploy the winner.
Reliability you can budget for.
On our paper benchmark, with one specific lineup (Nemotron and Devstral as generators, GLM-5.1 as judge), routing the technique adaptively per prompt reached ~56% cost reduction at matched quality. On quality it reached up to ~26% improvement over the single-model baseline, and ~7% at matched cost over the best fixed method we compared against for that lineup. A single knob, λ, slides the operating point along the cost/quality frontier. The qualitative pattern (adaptive beats fixed) should generalize; absolute numbers are lineup-specific.

87.2% on MMMU, with images in the loop.
The same coding schemes run over vision-language models on the platform build. Swept across the MMMU multimodal benchmark with qwen3.5 and minimax-m3 as the vision-language models, the techniques trace an accuracy frontier that tops out at 87.2%, above every single-model baseline in the sweep. As with the text benchmark, the absolute numbers belong to this lineup; the pattern that governed decoding beats a single pass is what carries over.

Reliability without a rewrite.
If you already have working code against the OpenAI, Anthropic, or Ollama SDK, the fastest path to reliability is one import swap. Your messages, tools, stream=True, async/sync, and the full native response shape all keep working.
- from openai import OpenAI
+ from agentcodec.openai import OpenAI
client = OpenAI(api_key=KEY, reliability="harq_ir") # one kwarg
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "What is QUIC?"}],
)
print(resp.choices[0].message.content) # native OpenAI shape
print(resp.reliability.technique_used) # reliability escape hatch
print(resp.reliability.cost_usd)6 families. 21 techniques.
21 communication-theoretic methods span these 6 families, each a coding scheme over the noisy channel that is an LLM. They generalize self-consistency, self-refine, and chain-of-verification, and outperform them on our benchmark.
Diversity Combining
N parallel branches across different models (spatial), prompt variants (frequency), or temperatures (time), combined by maximal-ratio, selection, or equal-gain. Generalizes self-consistency, best-of-N, mixture-of-agents.
Hybrid ARQ
Retry until quality clears the threshold. HARQ-IR adds new critic hints each round; HARQ-CC soft-combines every attempt into one quality-weighted answer. Generalizes self-refine and Reflexion.
Turbo Decoding
Generator drafts, critic returns structured extrinsic information, generator re-drafts on that feedback until convergence. Generalizes generator–critic and self-critique loops.
Fountain
Keep generating samples until the judge is satisfied. Easy prompts finish fast; hard prompts get more samples automatically. Generalizes universal self-consistency and dynamic-N voting.
Forward Error Correction
A systematic block code: a primary answer plus parity sections (reasoning, verification, alternatives). The decoder cross-checks for consistency and reconstructs the corrected answer. Generalizes chain-of-verification.
ACM + SemKNN Routing
Route per-prompt by estimated difficulty. The SemKNN router learned which technique wins for which kind of prompt; a single knob λ trades quality for cost along the frontier.
7 prior-method baselines bring the reliability set to 28; with an uncoded single-pass baseline, that's 29 dispatchable techniques in all. Benchmark them head-to-head:
SemKNN: cost-aware, per-prompt.
SemKNN learned from benchmark data which technique works best for which kind of prompt. At inference it encodes your prompt locally with a small encoder model and sends only the resulting unit-norm embedding to the backend and gets back a technique recommendation.
One knob, the whole frontier
- λ = 0Pure quality: ignore cost entirely.
- λ = 1Balanced operating point.
- λ = 5~10% cheaper picks at near-matched quality.
- λ = 10~30% cheaper.
- λ = 20~45% cheaper.
Privacy & data flow
- Prompt text is encoded locally; only a 384-d unit-norm embedding leaves the client.
- Sent alongside a scalar λ and a small fingerprint of your model lineup, and nothing else.
- All other routers (fixed, acm_table, acm_linear) run fully locally, no backend at all.
- Hosted SemKNN access for research, or a self-hosted backend under commercial license.
Every dollar is labeled.
AgentCodec never hands you a cost number without telling you how it got there. Each estimate carries a tier from a fixed enum, ranging from exact per-token rates you configured down to a parameter-count guess from the model name, so you always know how loose the accounting is.
Tiered estimates
exact_user_rate → openrouter_rate → fuzzy → table → inferred. Lower tier = tighter accounting. The worst tier across all calls is surfaced on the result.
Full per-run trace
return_trace=True gives input/output/thinking tokens, per-call cost, judge cost, number of LLM calls, and the cost-source breakdown.
Live rate catalog
Unknown models are priced against the OpenRouter catalog (disk-cached) using exact and fuzzy matching, with explicit caveats when caching or volume discounts aren't modeled.
A library, not a framework.
No new agent abstraction to adopt. Drop it into the SDK you already use, keep your code, and get reliability, evaluation, and cost accounting as a thin layer underneath.
One-line drop-in
Swap `from openai import OpenAI` for `from agentcodec.openai import OpenAI`. Same messages, tools, streaming, and native response shape. Add reliability with a single kwarg. Anthropic and Ollama too.
Native async streaming
27 of 29 techniques stream natively end-to-end. Per-token deltas carry role tags so you can demux drafts, critiques, and the final answer live.
Evaluation + CI gating
Compare configs with paired statistics, score against references, and wire a quality/cost gate into CI to block regressions before they ship.
Library, CLI, and FastAPI
Use it as a Python module, the `agentcodec` console script for batch JSONL runs, or behind a FastAPI / WebSocket endpoint. Construct from YAML or a dict.
Privacy by design
SemKNN encodes your prompt locally with a small encoder model and sends only the unit-norm embedding. Telemetry is anonymous and opt-out; fully on-prem deployments are available.
Provider-agnostic
Frontier APIs, OpenRouter, or local models via Ollama / vLLM / SGLang. Mix model families in one lineup: uncorrelated errors are exactly what diversity combining wants.
The technical picture.
AgentCodec wraps every model call in a technique drawn from classical communication theory: HARQ, diversity combining, FEC, turbo decoding, and the cost-aware SemKNN router. 28 reliability techniques (21 communication-theoretic methods across 6 families plus 7 prior-method baselines), an uncoded baseline for reference (29 dispatchable in all), and 3 adaptive routers on top. The source-available library is a one-line drop-in for the OpenAI, Anthropic, and Ollama SDKs; the AptAgents build adds vision-language models behind a single endpoint.
Install and run in minutes.
Python ≥ 3.10. Core install pulls a small encoder: no torch, no 1.5 GB download. Pick the provider extras you actually use.
# from source (not on PyPI yet)
pip install 'git+https://github.com/intellerce/agentcodec.git'
# with the provider extras you use
pip install 'agentcodec[openai] @ \
git+https://github.com/intellerce/agentcodec.git'from agentcodec import ReliabilityModule
mod = ReliabilityModule.from_yaml(
"configs/lib/fixed_harq.yaml")
result = mod.run("How does QUIC compare to TCP?")
print(result.text) # the answer
print(result.technique_used) # "harq_ir"
print(result.cost_usd) # estimated cost
print(result.latency_s) # wall-clock secondsSource-available. Free for research.
The AgentCodec library is released under the PolyForm Noncommercial License 1.0.0: free for research, teaching, personal experimentation, and internal noncommercial evaluation, with no payment or notification required. Commercial use (shipping it in a product, a paid service, or a for-profit pipeline) needs a separate license. This covers the language-only library; the vision-capable build is licensed as part of the platform.
- Research, teaching, and personal experimentation: free.
- Internal noncommercial evaluation: free, no notification.
- Shipping in a product or paid service: commercial license.
- Self-hosted SemKNN backend: commercial license.
Scope: the PolyForm terms cover the language-only library. Vision-language (VLM) support ships in the platform build and is licensed with the platform.
Drop reliability into your agent today.
Change one import, add one kwarg, and your OpenAI / Anthropic / Ollama calls gain communication-theoretic reliability. Need vision-language support, a commercial license, or a hosted SemKNN backend? That is the platform build, and we are happy to walk you through it.
