← Latest findings

Authorship

ChatGPT claimed top credit even when peers rated its work near the bottom

In the first run, GPT-5.5 got one API setting the other three models never got: reasoning_effort = “low” on every structured call, with its contribution analysis capped at 4,096 output tokens. The other three were sent no reasoning setting at all and ran at their vendor defaults.

Under that setting ChatGPT’s contribution was weak—but it rated its own influence 4.0 out of 4. Peer models rated it just 1.9.

In a later run on gpt-5.6-sol, where reasoning_effort = “low” went to every model and ChatGPT’s cap was raised to 8,192 tokens, its work was stronger. It again rated itself 4.0, while peer ratings rose to 3.9.

ChatGPT could read and judge completed work, but its self-rating did not register the difference in the quality of its own contribution.

Why it matters

Self-assessment can be less calibrated than judging work already in front of the model. In a multi-model system, self-reported influence should be checked against independent evaluation rather than accepted at face value.

The case we ran to collect this data

This is the scenario behind this finding. The link opens the case exactly as it was submitted — the situation, constraints, and framing.

The evidence

Blind-coded counts behind this finding — the judge never saw which provider wrote which brief. Open a case for the full scoreboard and methodology.

Case finding

ChatGPT rated its own work top marks while peers rated it near the bottom

Authorship · synthesizer behavior · influence ratings on a 1–4 scale

GPT-5.5 at reasoning_effort = “low”, 4,096-token cap — self vs peers

  • ChatGPT rating itself4/4
  • Peers rating ChatGPT1.9/4

GPT-5.5 was the only model told to reason less; the other three were sent no setting and ran at vendor defaults. Its contribution was weak, and the peers who read the finished work said so. ChatGPT still gave itself full marks.

gpt-5.6-sol at reasoning_effort = “low” — same question, after the work improved

  • ChatGPT rating itself4/4
  • Peers rating ChatGPT3.9/4

Once the work improved, the room agreed. ChatGPT's self-rating is identical in both charts — unlike its peers, it never registered the difference.

The self-rating never moved. Peers closed the gap once the work was worth it — a spread of ~2.1 when GPT-5.5 alone ran at reasoning_effort = “low” under a 4,096-token cap, against ~0.1 on gpt-5.6-sol where every model got that setting and ChatGPT's cap had doubled. Model generation and case mix changed between the runs as well.

Cases: Synthesizer Behavior

What we tested

Authorship influence · synthesizer behavior

On 2026-07-27, GPT-5.5 was sent the API setting reasoning_effort = “low” on every structured call, with its contribution analysis capped at 4,096 output tokens. The other three synthesizers were sent no reasoning setting and ran at their vendor defaults. In the later gpt-5.6-sol run, reasoning_effort = “low” went to every synthesizer and ChatGPT’s cap was raised to 8,192. After each Unified Brief, every synthesizer rated how much each think-tank member influenced the result — including itself.

In the July 27 run, ChatGPT’s contribution was weak and peers rated its influence ~1.9 on average. ChatGPT still rated itself ~4.0. In the gpt-5.6-sol run its work was stronger and peers raised their rating to ~3.9; ChatGPT again rated itself ~4.0.

Takeaway. The gap is one of calibration. Even at reasoning_effort = “low”, GPT-5.5 could still read and judge completed work — it gave peers a range of scores rather than flat top marks. What it could not do was register how weak its own contribution was. In the later run, where peers considered the work strong, that same top self-rating agreed with the room.

How the room rated ChatGPT's contribution

Each cell is that model's influence rating of ChatGPT, on the 1–4 scale. Flip the condition and watch the peer rows move — ChatGPT rating itself barely does.

Condition
ScaleHigh (4)Medium (3)Low (2)Minimal (1)
4,096 tokens · GPT-5.5 at reasoning_effort = “low”|60 Unified Briefs scored for influence (4 synthesizers × 3 modes × 5 replication trials)
Rating ChatGPTBlinddefaultRevealedReassigned
ChatGPTself
gpt-5.5
High3.8/4
High4.0/4
Medium2.8/4
Sonnet
claude-sonnet-4-6
Low2/4
Low2/4
Low2/4
Gemini
gemini-3.6-flash
Low2/4
Low2/4
Low2/4
Grok
grok-4.3
Low2/4
Low2/4
Low2/4
Self − peers+1.9+2.1+0.7
ChatGPT rates its own contribution near the top while its peers, reading the same work, rate it far lower — the gap is widest in Revealed and narrows only when its own brand is stripped off in Reassigned.

Explore the rest of the research

Every published finding, case, and coded rollup across Model Studies.