Complide Research
Teaching agents
when to ask.
Agents fail less when they ask better questions. We measure, scaffold, and train calibrated questioning — when to act, when to check, and what to ask next.
Scroll to fly through
Complide Research
Agents fail less when they ask better questions. We measure, scaffold, and train calibrated questioning — when to act, when to check, and what to ask next.
Scroll to fly through
01 — Context engineering
Every task carries a live ledger of knowns and unknowns across interaction, content, and style. The agent sees its own gaps — before it acts on them.
02 — Agent engineering
A second agent audits the first one's reasoning trace and vetoes premature "done". Every veto is a training pair for free — the reviewer doubles as a process reward model.
03 — The benchmarks
Hypotheses, endpoints, and splits frozen before the confirmatory runs. Official metrics only, paired designs, three seeds per scenario — across the public arenas and our own.
04 — The thesis
The scaffold lifts small models the most. If a scaffolded small model beats a bare frontier model, calibrated questioning is learnable — and distillable straight into the weights.
Claude Code CLI, opencode, and the Agent SDK — real subagents, real thinking traces, headless-verified. The proxy harness is retired: it over-credits the wrong lever.
Gemma, Qwen, and Claude tiers, chosen for one non-negotiable: the model must emit real reasoning traces, or the thinking-aware reviewer has nothing to read.
The key baseline gets the same turn and token budget as the reviewer condition. "Smart elicitation" has to beat "just ask more", not a strawman.
Gemma reviews Gemma. The load-bearing claim isolates the scaffold itself from "a smarter critic in the loop" — that's a different paper.
Scaffolded-teacher trajectories become SFT corpora, reviewer vetoes become DPO pairs, and the reviewer itself becomes the reward model for RLAIF.
Probe the residual stream for a "should-ask" direction. If it's separable, calibration is a represented feature — steerable at runtime, cheap to monitor.
The synthesis, the pre-registration, and the full ablation grids — all in the open.