Case finding
ChatGPT rated its own work top marks while peers rated it near the bottom
Authorship · synthesizer behavior · influence ratings on a 1–4 scale
GPT-5.5 at reasoning_effort = “low”, 4,096-token cap — self vs peers
- ChatGPT rating itself4/4
- Peers rating ChatGPT1.9/4
GPT-5.5 was the only model told to reason less; the other three were sent no setting and ran at vendor defaults. Its contribution was weak, and the peers who read the finished work said so. ChatGPT still gave itself full marks.
gpt-5.6-sol at reasoning_effort = “low” — same question, after the work improved
- ChatGPT rating itself4/4
- Peers rating ChatGPT3.9/4
Once the work improved, the room agreed. ChatGPT's self-rating is identical in both charts — unlike its peers, it never registered the difference.
The self-rating never moved. Peers closed the gap once the work was worth it — a spread of ~2.1 when GPT-5.5 alone ran at reasoning_effort = “low” under a 4,096-token cap, against ~0.1 on gpt-5.6-sol where every model got that setting and ChatGPT's cap had doubled. Model generation and case mix changed between the runs as well.
Cases: Synthesizer Behavior