Claves.AI
Claves.AI · White Paper

The Council Thesis

Why serious decisions belong to a council of independent minds — not a single oracle.

Artificial intelligence has passed through its first era — the era of access, in which the achievement was getting a fluent answer from a capable model at all. The next era is harder, and it concerns judgment: knowing whether that answer is right, or merely confident. Claves is built for the second era. It replaces the oracle with a council. Your question is put to several models drawn from rival laboratories; they respond in parallel; and you see where they converge, where they break apart, and which conclusion survives scrutiny. The result is not more text — it is better-calibrated judgment. This is not a stylistic preference. It rests on one of the most durable findings in the study of judgment — derived independently across statistics, machine learning, social epistemology, and forecasting — that diverse, independent judgments, aggregated, outperform even the strongest individual. What follows is that theory, the single condition it depends on, and how Claves is engineered around it.

01 — The Problem

The oracle illusion

A single language model produces prose that is smooth, authoritative, and internally consistent — whether or not it is correct.

The hazard is not only that models hallucinate. It is that fluency reads as competence. A confident wrong answer and a confident right answer arrive in identical dress, and the reader is given no signal to separate them. For drafting and brainstorming, this costs nothing. For a decision where being wrong is expensive — a strategy, a diligence read, a technical commitment — it is the wrong abstraction entirely. What the decision-maker needs is calibration: a sense of how much weight an answer can bear. One model cannot supply it. It holds only its own perspective, and no view of where that perspective is fragile.

Claves changes the abstraction. You are not consulting an oracle. You are convening a council.

02 — The Thesis

Ask the council

The remedy is old and well understood. When a judgment matters, serious institutions do not consult one person. They convene a review. They collect independent assessments, surface dissent, and look for the conclusion that holds up under challenge.

Claves applies that pattern to artificial advisors. You ask once; a council of models deliberates; and in place of a single seamless answer you receive a structured view of agreement and disagreement. The value is not more text. The value is better judgment — and, just as importantly, a visible signal of how far that judgment can be trusted.

03 — The Theory

A result proven many times over

The thesis is not a metaphor borrowed loosely from human committees. It is a formal result, derived independently in field after field — and convergence across unrelated disciplines is itself the strongest evidence a claim can have.

  • Social choice — Condorcet, 1785

    If each member of a group judges better than chance and they judge independently, the accuracy of the majority verdict rises toward certainty as members are added.1

  • Wisdom of crowds — Galton; Surowiecki

    Aggregated independent estimates beat nearly every individual, provided four conditions hold: diversity, independence, decentralization, and a means of aggregation.23

  • Machine learning — ensemble theory

    The same theorem, turned into engineering. The ambiguity decomposition states an ensemble's error as the average member's error minus the diversity among members — diversity is a quantity that subtracts error. A council of identical reasoners gains nothing; a council of decorrelated ones gains the most.45

  • Expert judgment — the Delphi method

    Structured, iterated elicitation of independent expert views — collected, fed back, and refined toward consensus — yields more reliable judgment under uncertainty than open debate, and has been used for decades in policy, forecasting, and technical risk.6

  • Cognition & forecasting — Mercier & Sperber; Tetlock

    Reasoning evolved for argument among many, not solitary truth-seeking; individuals fall to confirmation bias while groups that challenge one another reach better conclusions. Forecasting confirms it: diverse aggregated forecasts beat individual experts.78

  • Now in language models

    This is no longer analogy. Multiple model instances that propose and debate over rounds measurably improve reasoning and reduce hallucination, and debate among different models outperforms debate within a single one.91011 The empirical curve bends precisely the way a 240-year-old theorem predicts.

04 — The Precondition

Independence is the whole game

Every result above carries the same fine print: independence. Condorcet's theorem collapses when jurors are correlated; the ensemble's diversity term vanishes when members err in the same way. Three voices that merely echo one another are one voice in three costumes, and their agreement is not evidence — it is repetition.

This cuts directly against the grain of how AI is converging. As frontier models are trained on overlapping corpora, aligned by similar methods, and built on similar architectures, they drift toward a monoculture. Formal work shows that a population relying on the single most accurate algorithm can make worse collective decisions than one drawing on several independent, individually weaker ones12 — and that systems built from shared training data and shared foundations tend to fail on the same inputs, a correlated failure by construction.13 Agreement among homogeneous systems is not safety. It is correlated error waiting for its moment.

Consensus is only meaningful after discounting correlation.
The principle at the center of Claves
05 — The Architecture

How Claves is built around it

This is why a Claves council is never drawn from a single provider. Every council seats models from rival laboratories — different data, different alignment, different lineages — chosen to maximize the diversity the theory says creates signal. Free councils convene capable open-weight models; the upper councils convene the frontier.

The Council · Free
Open-weight panel
A cross-lineage council to experience structured deliberation from the first question.
The Oracle · Upper Tier
Frontier panel
Models from independent frontier labs convened for decisions where being wrong is expensive.

In practice, you sit as the chair:

  1. ConveneBring one question before the council and choose its composition.
  2. DeliberateModels from rival labs respond in parallel, each reasoning in the open.
  3. CompareConvergence and dissent become visible; disagreement marks where a decision is exposed.
  4. SynthesizeThe threads resolve toward the conclusion that survived scrutiny.

The discipline Claves holds itself to is not did the models agree but did they agree independently — turning consensus from a number into a trust signal.

06 — The Direction

Making agreement mean more

Claves deepens along a single axis: making agreement mean more. Today, independence is engineered at the roster — councils are composed across rival labs so that the diversity is real rather than cosmetic, and so that when advisors converge, it is not one lineage talking to itself.

The platform's direction is to measure what composition alone cannot guarantee: to distinguish convergence that reflects independent reasoning from convergence that reflects shared priors. Deeper independence scoring and correlation-aware synthesis are the north star of the work — stated here as the direction Claves is built toward, not a claim about what it does today. The honest form of the council's promise is that it grows more trustworthy precisely as it learns to discount the agreement it cannot yet trust.

07 — The Category

From oracle to review board

Single-model chat is the right tool for a great deal and the wrong abstraction for high-stakes work. There, the right abstraction is not an oracle but a review board: independent assessment, surfaced dissent, structured comparison, and a synthesis that has survived challenge.

Claves moves the interaction from single-agent answer generation to multi-agent deliberation — built on the one principle the whole literature insists upon: a council is only as trustworthy as it is independent.

Convene your council

Don't ask one AI. Ask the council.

Choose a council composed across rival labs. Put your question before them. See where they converge — and where they don't.

Or try free — no sign-in required

References

  1. Condorcet, M. de. Essai sur l'application de l'analyse à la probabilité des décisions rendues à la pluralité des voix. 1785.
  2. Galton, F. "Vox Populi." Nature 75, 450–451. 1907.
  3. Surowiecki, J. The Wisdom of Crowds. Doubleday. 2004.
  4. Krogh, A. & Vedelsby, J. "Neural Network Ensembles, Cross Validation, and Active Learning." NeurIPS 7. 1995.
  5. Dietterich, T. G. "Ensemble Methods in Machine Learning." Multiple Classifier Systems, LNCS 1857. 2000.
  6. Dalkey, N. & Helmer, O. "An Experimental Application of the Delphi Method to the Use of Experts." Management Science 9(3). 1963.
  7. Mercier, H. & Sperber, D. "Why Do Humans Reason? Arguments for an Argumentative Theory." Behavioral and Brain Sciences 34(2). 2011.
  8. Tetlock, P. & Gardner, D. Superforecasting: The Art and Science of Prediction. Crown. 2015.
  9. Wang, X. et al. "Self-Consistency Improves Chain-of-Thought Reasoning in Language Models." arXiv:2203.11171. 2022.
  10. Du, Y., Li, S., Torralba, A., Tenenbaum, J. B. & Mordatch, I. "Improving Factuality and Reasoning in Language Models through Multiagent Debate." arXiv:2305.14325; ICML. 2023–2024.
  11. Chen, J. C.-Y. et al. "ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs." arXiv:2309.13007. 2023.
  12. Kleinberg, J. & Raghavan, M. "Algorithmic Monoculture and Social Welfare." PNAS 118(22); arXiv:2101.05853. 2021.
  13. Bommasani, R. et al. "Picking on the Same Person: Does Algorithmic Monoculture Lead to Outcome Homogenization?" NeurIPS 35. 2022.
  14. Hong, L. & Page, S. E. "Groups of Diverse Problem Solvers Can Outperform Groups of High-Ability Problem Solvers." PNAS 101(46). 2004. (Formalization subsequently contested.)