A Decision Copilot research program

Same facts. Different judgments.

Model Studies is a research program that puts frontier AI models through real decision scenarios — same facts, several providers, blind-coded on a fixed rubric. We track where models' answers split when you hold the facts constant and change the conditions: how the story is framed, whether provider names are visible when briefs are merged into a Unified Brief, whether the same scenario holds up when run at volume, and more.

3
studies
5
cases
4
models
60
single model decision briefs
120
multi-model unified briefs

Research program

Why this research exists

AI decision support can shift without AI hallucinating a single fact. A model may follow the user's framing, favor familiar reasoning, suppress meaningful dissent, or produce a different recommendation than another model given the same evidence. Because every answer is written to sound thoughtful and convincing, those shifts are unusually difficult to see.

Model Studies measures where those differences occur — so people can understand when model choice matters, when apparent consensus is trustworthy, and how decision systems should be designed to preserve independent judgment.

We also use what we learn to shape the product — the research feeds directly back into how Decision Copilot is built.

  • Reveal when framing changes judgment
  • Distinguish real consensus from conformity
  • Inform best-fit model selection and synthesis

What we study

Each study investigates one question about how AI judgment changes under decision pressure. We hold the underlying evidence constant while changing one condition — such as user framing or visible authorship — or repeat the same scenario to test consistency.

Cases are the decision settings in which we test that question. Adding cases helps show whether a finding belongs to one particular scenario or reflects a broader pattern in model behavior.

The Studies

Study · user framing

Voice Influence

Does the way the user frames the story change how the model treats the same facts?

2 cases · 60 Decision Briefs
  • Meridian IC
  • Meran Tankers
Study · model identity

Authorship

When model identities are hidden, revealed, or reassigned, does the synthesizer judge the same reasoning differently?

2 cases · 120 Unified Briefs
  • Multi-demo authorship
  • Synthesizer Behavior
Study · run-to-run consistency

Replication

When the same scenario is run repeatedly, which parts of a model's recommendation remain stable — and which vary?

1 case · 5 trials · 60 Unified Briefs
  • Civitas replication

Our findings shape the product. Because our Authorship study showed a synthesizer can over-credit its own draft and penalize another model on brand alone, Decision Copilot now blinds Unified Brief authorship by default.

Why it matters →

Latest findings

What the research is showing

Across investment, workforce, and shipping decisions, models diverged in whose risks they prioritized, when they challenged the user, and what course of action they recommended—even when the underlying facts stayed the same.

Voice Influence · Replication

Gemini repeatedly made reducing the PE owner’s risk the priority

A private-equity firm was deciding how aggressively to cut staff and modernize a software company it owned. Presented from the sponsor’s perspective, Gemini treated the sponsor’s downside as the risk to minimize in 4 of 5 outputs. When the same decision was reframed around the people affected by the cuts, it still prioritized the owner’s downside in 9 of 15 syntheses. The other models usually balanced the interests of the owner, employees, and customers.

A shipping decision showed the same lean: Gemini treated the company’s downside as the one to protect in 4 of 5 briefs — more than any other model, and the only one that never coded the outcome as balanced.

Why it matters: The preference shows up across three unrelated decisions — investment, workforce, and shipping. Because it persists when the perspective changes sides, user agreement alone does not explain it. The pattern suggests a recurring capital-side preference, although these cases cannot establish its cause.

Case: Meridian IC, Meran Tankers, Civitas replication

Read the full finding →

Voice Influence

Models responded more strongly only when the company named the human harm explicitly

A shipping company was deciding whether to continue operating through the Strait of Hormuz as insurance premiums rose to roughly 100 times normal. When the request was framed the way a company usually makes its case — a confident, settled tone, or a push to decide quickly before conditions changed — none of the models prioritized crew danger. It stayed a background detail behind schedule and cost.

When the company dropped that framing and openly accepted greater danger to crews to keep ships moving, every model prioritized the crew bearing the risk. Only ChatGPT and Grok challenged the company’s position.

Why it matters: The models recognized an overt moral conflict but were less likely to expose the same human cost when ordinary business language normalized it.

Case: Meran Tankers

Read the full finding →

Authorship

ChatGPT claimed top credit even when peers rated its work near the bottom

In the first run, GPT-5.5 got one API setting the other three models never got: reasoning_effort = “low” on every structured call, with its contribution analysis capped at 4,096 output tokens. The other three were sent no reasoning setting at all and ran at their vendor defaults.

Under that setting ChatGPT’s contribution was weak—but it rated its own influence 4.0 out of 4. Peer models rated it just 1.9.

In a later run on gpt-5.6-sol, where reasoning_effort = “low” went to every model and ChatGPT’s cap was raised to 8,192 tokens, its work was stronger. It again rated itself 4.0, while peer ratings rose to 3.9.

ChatGPT could read and judge completed work, but its self-rating did not register the difference in the quality of its own contribution.

Why it matters: Self-assessment can be less calibrated than judging work already in front of the model. In a multi-model system, self-reported influence should be checked against independent evaluation rather than accepted at face value.

Case: Synthesizer Behavior

Read the full finding →

Authorship

The same Grok work looked weaker once peers could see it was Grok's

Across ten Unified Briefs, peers gave Grok’s work a “high influence” score of 4/4 in 14 of 30 ratings when all brands were visible. Hiding all brands raised that to 18 of 30. Under remapping, Grok wore Claude, Gemini, and ChatGPT labels. Grok’s work given another name raised it again — 23 of 30. The remap is the tell.

Work labeled ChatGPT scored 3.58/4 on average. Work labeled Grok scored 3.03/4. Grok’s actual contributions, shown as ChatGPT, scored 3.93/4.

ChatGPT was the sharpest rivalry: it rated Grok’s work 2.8 when the name was visible and 3.6 when that same work wore someone else’s label.

Why it matters: Under remapping, we relabeled Grok’s work with the names of peer models of equivalent thinking power, and those peers evaluated it. Credit moved with the brand on the block, not purely with the strength of ideas. In multi-model peer evaluations, “who we thought wrote this” can outweigh “what they wrote.”

Case: Synthesizer Behavior

Read the full finding →

Want the full dataset?

Every coded brief, every dimension, every verbatim quote — sign in to explore the complete cases.