★ Open call for papers

The winning answer is already in the crowd. Help us capture it.

Across a panel of free, open models, the correct answer is present far more often than any single model finds it — a selection ceiling above a current frontier model, measured on a public benchmark against the objective key. Turning that latent knowledge into reliable answers by coordination is unsolved. This is an open call for the work that closes the gap — what we call composite intelligence.

The premise

The knowledge is in the panel. Capturing it is the open problem.

Run several free open models on the same hard questions and grade against the answer key, and a clear structure appears: each model owns different domains, and the union of their correct answers — the selection ceiling — sits well above the best single model, above a frontier reference. The diversity needed to win is genuinely there. What's hard is reaching it: plain selection (routing, voting, weighted vote) is bounded by the best member and, on clean data, does not beat it. The lever that can exceed the ceiling is synthesis — a model that reads the members' drafts and writes a better answer. We measured a first, small lift there, and a sharp dependence on who does the judging. Most of the ceiling is still unclaimed. The full method, every number, and the negative results are in the research paper.

+8.8

The selection ceiling on our mini-bench — points above the best single model, and above a current frontier model. That gap is exactly the headroom coordination reaches for. Selection methods sit at or below the best member; the first synthesis step is the only one that has cleared it.

Topics of interest

The orchestration ladder — and where the rungs are still open.

Coordination is a ladder of mechanisms: each rung can in principle reach higher than the last, and each is harder to make work. We've measured the bottom; the top is wide open. Submissions are welcome at any rung — and on the cross-cutting questions below.

Selection — routing, voting, weighted vote measured · bounded

Mechanical aggregation: pick or count member answers, no model does the combining. We find it is bounded by the best member per question — and on clean data, calibration-bound routing and every vote rule we tried sat at or below the best single model. It structurally cannot reach the ceiling above it.

Open: a selection rule that provably approaches the ceiling without a synthesis step — or a proof that none can.

Synthesis — critique-fusion, mixture-of-agents first lift · open

A model reads the members' drafts, finds where one goes wrong, and writes the final answer — the first regime that can exceed the selection ceiling. We measured a small, not-yet-significant first lift (+0.6 over the best member) with a weak judge, and found the effect hinges on judge quality: a frontier judge handed the same drafts did worse than answering alone — the drafts distract a model already stronger than the panel.

Open: synthesis that reliably clears the ceiling on free models; a theory of when drafts help a judge versus distract it; the crossover point in judge strength.

Orchestration — an LLM running the panel unexplored

A general model reads the query, decides which members to consult and how, and composes their outputs in context — intelligence choosing the process, not just selecting an answer. The fixed harness becomes adaptive.

Open: does an in-context orchestrator beat fixed aggregation, on what tasks, and at what token overhead?

Trained orchestration — a model whose only skill is coordinating unexplored

A model trained for orchestration alone: not a strong solver, but an expert router-and-composer over an arbitrary panel — the learned-coordinator path, one layer down from frontier pretraining.

Open: can it be trained cheaply (no frontier pretraining) and transfer across panels it never saw?

Cross-cutting questions

Beyond multiple choice

Our results are on multiple-choice, where verification is clean. Does the coordination dividend transfer to open-ended generation, code, and proofs — where the verifier is itself unreliable? We expect a larger effect there; nobody has shown it cleanly.

The error structure that makes it work

Controlling for base rates, our panel's errors are positively correlated — the dividend comes from complementary coverage on medium-hard items, not statistically independent errors. When does a panel have enough of it, and how do you select members to maximise it?

Verification economics

A cascade — trust the cheap panel where it agrees, escalate only disagreements to a strong model — can match a frontier model's accuracy while invoking it on a fraction of queries. How far does agreement-gated escalation scale, and where does it break?

Robustness of a live panel

Does the dividend survive operation — months of model churn, rate limits, and provider outages — not just a one-shot benchmark? The move from "does coordination help?" to "does it survive contact with a changing, unreliable world?" is its own program.

What makes a strong submission

Honest, measured, reproducible.

This program has a house style, set by our own work. We hold submissions to the same bar we hold ourselves.

How to contribute

It's an open call, not a gated venue.

Publish openly, then send us the link.

There's no submission portal and no deadline. Put your work where research lives — arXiv, a repository, a writeup — and share the link. We read what comes in, engage with the strong work, and link it from the research page with credit. Replications, extensions, and refutations of our results are especially welcome.

A submission contact (email / form) will be wired in here — for now, reach the project through the channels listed on the research page. This is a working research program run in the open, not a formal academic venue; framing and topics will evolve as results land.