Advancing Technology
Advancing Technology
Council of experts Illustration

The Appearance of Knowing

Listen to the article

In 1950, a young economist gave himself a problem that looked almost too simple to be serious. Kenneth Arrow wanted to know whether a group of people could, together, reach a choice that every one of them would consider fair. Suppose the members of a committee, or the citizens of an entire country, have to choose from several options, and that each person ranks those options by preference, from the one they want most to the one they want least. Is there a fair way to combine all of those individual preferences into a single decision the group can accept? Europe and Asia had just fought a war over how people should be governed, so the question was not abstract. Almost everyone assumed a fair rule existed; the only work left was to find it.

Arrow found something else. He proved that no such rule can exist. Not that no one had found it yet, but that it is impossible: any method that respects a few plain conditions of fairness will, in some situations, contradict itself. The proof was solid enough to found a new field, social choice theory, and twenty years later it helped earn him a Nobel prize. A simple question about voting had a permanent and unwelcome answer.

Arrow was asking about people, but the question reaches further. Can many independent judgments be combined into a single decision that is fair to all of them? He asked it of voters; we now ask it of LLMs. When an answer from a language model leaves us unsure, a reflex could be to gather more voices around it, to give the model an expert role, to convene a panel, to build a council of models that argue and then agree, on the assumption that more voices mean more rigor. Arrow’s old problem is waiting for us at the end of that road, simply transposed. But all of this begins more modestly, with two magic words.

The magic words

It began about a year ago, which in this field already feels distant. The advice was everywhere: to get a better answer from a language model, open with the magic words, “You are an expert in…” Name the role, and the quality would follow. It was free, it was simple, and it felt like a key that opens a lock.

The model makers recommended it themselves. Google, OpenAI, and Anthropic all placed “You are an expert in…” templates in their own guides, and the rest of us copied them.¹ The logic sounded right: a model trained on human writing imitates the role it is given. The proof that this changed the answers, and not only their tone, was, in fact, quite weak from the start.

When researchers tested the trick on questions with measurable answers, the expert role did not reliably help. On factual and reasoning tasks it sometimes made things slightly worse.² The costume changes the voice, not the knowledge. Sounding right is cheap, but being right is not.

The expert role, seen plainly, is one model and one answer made to look like more than they actually are. It meets Arrow from the other side. We keep trying to turn one voice into many. He was trying to turn many voices into one.
Both reach the same limit.

What real looks like

If the role is only a costume, the real question is what genuine expertise looks like underneath. A team at Stanford built one answer, and what people do with it afterward is the real lesson. They called it STORM, presented at NAACL in 2024.³ The name is an acronym, and the important word sits in the middle: the ‘R’ stands for Retrieval.

Before writing, the system researches: it searches the internet, collects real references, and lets imagined writers with different viewpoints question an expert who answers from those sources. Then it composes the article, with citations. The personas are not the engine; they are a way to ask sharper questions. The answers come from documents the system actually found. Against a strong retrieval baseline, its articles were judged clearly better organized and broader in coverage.³ The researchers were honest about the limits: bias inherited from sources, and also unrelated facts wrongly connected together.³

This gives us a lens for every later case. When a method presents itself as rigorous, the question to ask is simple. Where is the independence?
STORM has one kind: its information is independent, grounded in real and varied sources rather than produced from inside the model itself. That is what earns the gain.

This leads us to the shortcut that’s worth rejecting, one that is popular on social networks. Keep the five experts, drop the sources. “Simulate five expert perspectives,” with nothing real attached, is not five experts.
It is one model writing five plausible answers in five voices, then summarizing itself. They share one mind, one training set, one set of gaps. The shortcut borrows STORM’s measured gain and removes the part that earned it. The expert role gave the appearance of authority; the panel gives the appearance of debate. Neither adds a single fact.

Degrees of independence

So one independence comes from real sources. The other comes from the models themselves: different providers, different training data, so they do not tend to make the same mistakes. The two are easy to confuse and worth keeping apart. The first grounds the answer in the world. The second lets the voices fail for different reasons.

The options run from least independent to most. Weakest is the single expert role. Then the imagined panel, one model playing five parts. A real improvement is a council drawn from one provider: many agents, well organized, even told to disagree, yet all running on one company’s model, so the same gaps sit beneath every voice. At the far end is a council of different providers, trained on different data, ideally reading real sources. Only there do agreement and disagreement carry information.
Andrej Karpathy’s LLM Council ⁴ does exactly this: models from several companies answer, review each other anonymously so no model favors its own answer, and a final model writes the result. It is also, not by accident, the most expensive.

I have built one of these myself. My council was not a crowd of identical voices but a small panel of different lenses, each given its own task, each told to challenge the others rather than agree, each made to read the real material before deciding. It worked. It caught the kind of quiet mistake a single reading misses: a check that proves nothing, a change that looks safe and is not. The harder lesson took longer. My panel disagreed in useful ways, but every voice ran on models from one company (A\ – not to mention). It had one independence and not the other: it read real sources, yet it could not think with a second mind. Different roles are not completely different models. The one mistake such a council cannot catch is the mistake its own base model would also make.

Matching the tool to the need

A real council, then, buys a real but limited gain. It can catch the flaw a single answer would miss. The limit is what the research keeps finding: the improvement appears early, then stops growing after a few rounds , and in careful comparisons, these systems do not reliably beat much simpler methods.⁵

The simplest is easy to describe. Ask one model the same question several times, since its answers vary from one run to the next, then keep the answer it returns most often. Surprisingly enough, even against that, a council hardly shows any reliable advantage.

This is easy to misread what was just said as a defense of the magic words. It is not. The single model that matches a council is never the one given a bare role. It is one made to work: the question split into parts, real sources placed before it, its first answer sent back for criticism. Only “You are an expert in…” costs nothing, and only it buys nothing.

Which gives us a rule and a short menu. Match the tool to the need, not to the most elaborate option.

1) For grounded facts and citations, the answer is retrieval: STORM, free and open, or any model that is given real sources to read.

2) To test a question of judgment, ask a few different providers’ models (Gemini, GPT, Claude,…) the same question and read where they disagree.

3) For fast, structured disagreement, a single-provider panel of distinct roles works well, as long as we remember that its agreement is not proof.

Underneath all three, one principle: the value never comes from the number of voices. It comes from adding something independent, real sources or a second mind.

Counting the votes

Suppose we do everything right. Different providers, real sources, true independence between the voices. One problem still waits at the end, and it is Arrow’s. Someone has to decide what all those voices concluded.

There are two doors, and both disappoint. Send the whole exchange to a final model to summarize, and the independence we paid for collapses back into one mind at the last step. Hand it to a person, and we ask one reader to hold more analysis than any single person comfortably can.

Arrow studied people voting in elections. He never saw a machine answer a question. And yet the theorem his voting problem produced now reaches our councils too, if only by analogy. We have taught one judge to consult many others. Whether that frees it from being a single judge, or only teaches it to speak in many voices, is the question he handed us seventy years before we thought to ask it.


References

  1. “Prompting Science Report 4: Playing Pretend: Expert Personas Don’t Improve Factual Accuracy.” 2025, arXiv:2512.05858. Documents that Google (Vertex AI), Anthropic, and OpenAI recommend persona prompting in their official guides. On the origin of the pattern, see White, J. et al., “A Prompt Pattern Catalog to Enhance Prompt Engineering with ChatGPT,” 2023, arXiv:2302.11382. https://arxiv.org/abs/2512.05858
  2. Zheng, M., Pei, J., Logeswaran, L., Lee, M., & Jurgens, D. “When ‘A Helpful Assistant’ Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models.” 2024, arXiv:2311.10054. https://arxiv.org/abs/2311.10054
  3. Shao, Y., Jiang, Y., Kanell, T., Xu, P., Khattab, O., & Lam, M. “Assisting in Writing Wikipedia-like Articles From Scratch with Large Language Models.” Proceedings of NAACL 2024, pp. 6252-6278. https://aclanthology.org/2024.naacl-long.347/
  4. Karpathy, A. “LLM Council.” Open-source project, 2025. https://github.com/karpathy/llm-council
  5. “Debate or Vote: Which Yields Better Decisions in Multi-Agent Large Language Models?” 2025, arXiv:2508.17536, and “Literature Review of Multi-Agent Debate for Problem-Solving,” 2025, arXiv:2506.00066. https://arxiv.org/abs/2508.17536
  6. Arrow, K. J. Social Choice and Individual Values. Yale University Press, 1951.

Stay in the loop

New essays on AI, technology, and society, delivered when they matter.

Powered by Buttondown

Serendipity illustration 1

The Missed Discoveries

Prev