A miniature observatory pointing its telescope at a question-mark constellation
A great ledger book of knowns and unknowns tended by scribe robots
A reviewer robot reading another robot's stream of thought through a lens
A miniature benchmark arena with robots working through scored lanes
A chart-shaped mountain a small robot climbs using a glowing scaffold

Complide Research

Teaching agents
when to ask.

Agents fail less when they ask better questions. We measure, scaffold, and train calibrated questioning — when to act, when to check, and what to ask next.

Scroll to fly through

01 — Context engineering

An honest ledger
of unknowns.

Every task carries a live ledger of knowns and unknowns across interaction, content, and style. The agent sees its own gaps — before it acts on them.

Uncertainty ledger Slot checklist EVPI-calibrated

02 — Agent engineering

A reviewer that
reads the mind.

A second agent audits the first one's reasoning trace and vetoes premature "done". Every veto is a training pair for free — the reviewer doubles as a process reward model.

Thinking-aware Separation of duties Veto → DPO pair

03 — The benchmarks

Pre-registered, or it
didn't happen.

Hypotheses, endpoints, and splits frozen before the confirmatory runs. Official metrics only, paired designs, three seeds per scenario — across the public arenas and our own.

ReqElicitGym NoisyToolBench AgentDojo

04 — The thesis

Context engineering
buys scale.

The scaffold lifts small models the most. If a scaffolded small model beats a bare frontier model, calibrated questioning is learnable — and distillable straight into the weights.

Benefit ↑ as size ↓ Data factory DPO · SFT · RLAIF
4axes of iteration: harness · model · benchmark · technique
0.13 → 0.39elicitation lift from the inference-time scaffold
H1–H3hypotheses pre-registered before confirmatory runs
Freetraining pairs — every reviewer veto is one

The program,
end to end.

Harness lab

Claude Code CLI, opencode, and the Agent SDK — real subagents, real thinking traces, headless-verified. The proxy harness is retired: it over-credits the wrong lever.

Model ladder

Gemma, Qwen, and Claude tiers, chosen for one non-negotiable: the model must emit real reasoning traces, or the thinking-aware reviewer has nothing to read.

Compute-matched controls

The key baseline gets the same turn and token budget as the reviewer condition. "Smart elicitation" has to beat "just ask more", not a strawman.

Same-model reviewer

Gemma reviews Gemma. The load-bearing claim isolates the scaffold itself from "a smarter critic in the loop" — that's a different paper.

Data factory

Scaffolded-teacher trajectories become SFT corpora, reviewer vetoes become DPO pairs, and the reviewer itself becomes the reward model for RLAIF.

White-box probes

Probe the residual stream for a "should-ask" direction. If it's separable, calibration is a represented feature — steerable at runtime, cheap to monitor.

Read it before you believe it.

The synthesis, the pre-registration, and the full ablation grids — all in the open.