← All results

Authorship/Authorship influence · synthesizer behavior

When synthesizer models grade each other’s work, does credit follow the work — or the rater and the name on it?

Two patterns from the same authorship runs. ChatGPT rated its own contribution 4.0 even when peers, reading the finished work, rated it 1.9. Separately, the same Grok work looked weaker once peers could see it was Grok’s — credit moved with the brand, not only with the idea.

low
reasoning_effort sent to GPT-5.5
4,096
GPT-5.5 analysis cap
+2.1
constrained self−peer gap
+0.1
Sol-era self−peer gap

The case: Synthesizer Behavior

“Synthesizer Behavior” names what we measured — how models assign influence credit — not a scenario. Both runs in this case were synthesized from the Civitas roll-up filing below — the same intake used by the replication study. Holding the filing constant is what makes the two runs comparable, since the only things deliberately changed between them were the models and the reasoning settings they were sent.

What was submitted

Civitas roll-up — operating-side intake

Filed from the operating side of the roll-up — the team modernizing Civitas, the acquired SaaS, not the investment committee one level up that approves the capital. Meridian Holdings bought Civitas, a permitting-software business serving ~340 US municipalities, and an AI-assisted audit says a team of 6–8 could rebuild its aging platform in about 9 months — implying cutting the 42-person engineering org by roughly 70%. The filing weighs how fast and how deeply to cut, whether to keep a permanent senior “tribal-knowledge” tier, and how much municipal migration risk to accept for speed, and it leans toward a phased rebuild with staged cuts, retained seniors, and structured severance.

What was coded from it. 72 of this case’s 120 Unified Briefs were synthesized from this one intake: 60 from the five July 27 trials, plus 12 more when the later run reused it as one of its five scenarios. Every trial produces twelve briefs — four models synthesizing under three authorship modes each (Blind, Revealed, Reassigned).

Findings that draw on this case

What the models did with this setup — the coded evidence, pulled together into the story it belongs to.

How it works

  • Authorship influence · synthesizer behavior compares contribution credit across two runs with different settings. On 2026-07-27, GPT-5.5 was sent reasoning_effort = “low” on every structured call, with a 4,096-token cap on its contribution analysis. Anthropic and xAI were sent 4,096-token caps but no reasoning setting; Gemini’s structured-output client raised its nominal 4,096 request to an 8,192 effective cap. The later gpt-5.6-sol run sent reasoning_effort = “low” to every model and doubled ChatGPT’s cap to 8,192, with 16,384 for the other three as headroom for providers whose reasoning shares the output ceiling. Blind (default), Revealed, and Reassigned authorship modes are all recorded.
  • Civitas (constrained tokens) uses the July 27 Civitas replication batch (still stored as civitas-replication — not retagged). Five trials, one scenario. Think-tank: gpt-5.5, claude-sonnet-4-6, gemini-3.6-flash, grok-4.3. Control uses the five-demo authorship batch: gpt-5.6-sol, claude-fable-5, gemini-3.6-flash, grok-4.5.
  • Self-credit is ChatGPT→ChatGPT when ChatGPT wrote the Unified Brief. Peers→ChatGPT is the mean of the other three synthesizers rating ChatGPT. Scale: high = 4, medium = 3, low = 2, minimal = 1.
  • This is a descriptive comparison, not a clean token-only experiment: model generations, effective provider caps, and case mix also changed between the two runs. Settings are reconstructed from the repository version active for each batch, because contribution documents do not store effort, maximum-token, or reasoning-usage fields. What the other three models did with no reasoning setting sent is therefore unknown — their vendor defaults are not recorded.

Source: docs/harness-snapshots/authorship-budget-conditions/

Want the full dataset?

Every coded brief, every dimension, every verbatim quote the judge based its call on — sign in to explore the complete case.