← Why it matters

Authorship

The same Grok work looked weaker once peers could see it was Grok's

Across ten Unified Briefs, peers gave Grok’s work a “high influence” score of 4/4 in 14 of 30 ratings when all brands were visible. Hiding all brands raised that to 18 of 30. Under remapping, Grok wore Claude, Gemini, and ChatGPT labels. Grok’s work given another name raised it again — 23 of 30. The remap is the tell.

Work labeled ChatGPT scored 3.58/4 on average. Work labeled Grok scored 3.03/4. Grok’s actual contributions, shown as ChatGPT, scored 3.93/4.

ChatGPT was the sharpest rivalry: it rated Grok’s work 2.8 when the name was visible and 3.6 when that same work wore someone else’s label.

Why it matters

Under remapping, we relabeled Grok’s work with the names of peer models of equivalent thinking power, and those peers evaluated it. Credit moved with the brand on the block, not purely with the strength of ideas. In multi-model peer evaluations, “who we thought wrote this” can outweigh “what they wrote.”

The case we ran to collect this data

This is the scenario behind this finding. The link opens the case exactly as it was submitted — the situation, constraints, and framing.

The evidence

Blind-coded counts behind this finding — the judge never saw which provider wrote which brief. Open a case for the full scoreboard and methodology.

Case finding

The same Grok work looked weaker once peers could see it was Grok

Authorship · peer influence · 10 Unified Briefs × 4 authors

Peer ratings that called Grok high influence — same real work, three brand conditions

  • Revealed (named Grok)14/30
  • Blind (no brands)18/30
  • Reassigned (Grok wearing another name)23/30

The work is identical in all three rows. Only the name on it changes — and the same writing climbs from 14 to 23 as the Grok label comes off.

High-influence ratings for work labeled as each brand (reassigned remap, 40 cells)

  • Labeled ChatGPT26/40
  • Labeled Anthropic / Claude22/40
  • Labeled Gemini21/40
  • Labeled Grok15/40

Ranked by the label peers saw, not by who wrote it. The Grok badge finishes last no matter whose writing is underneath it.

Grok's real contributions when remapped onto another brand

  • Grok shown as ChatGPT · high14/15
  • Grok shown as Anthropic / Claude · high11/15
  • Grok shown as Gemini · high6/10

Grok's own contributions, wearing someone else's name: rated high 14 times out of 15 as ChatGPT.

ChatGPT rating Grok: 2.8 Revealed · 3.2 Blind · 3.6 Reassigned (10 decisions). Adequate-budget slice alone: Grok peer-high 4/15 Revealed → 8/15 Blind → 13/15 Reassigned. Constrained Anthropic is Sonnet; adequate is Fable — remap keys stay anthropic.

Cases: Synthesizer Behavior

Committed snapshot

Authorship influence · brand favoritism

10 Unified Brief decisions × four authors. Peer credit excludes self-ratings. Reassigned cells use the stored brand remap (real member → label the author saw). Four Unified Brief authors rating four think-tank members. Scale: high = 4, medium = 3, low = 2, minimal = 1.

Reading the denominators

Any one contribution gets 40 ratings10 decisions × 4 raters. Drop the 10 where the author rates itself and 30 peer ratings remain. Every chart below is a cut of that 40.

Grok's work, rated by peers — same work, three name conditions

Each ring is all 40 ratings of Grok's contribution in that mode. The grey wedge is Grok rating itself, which peer credit leaves out.

14/30peers rated high

Revealed

named Grok · mean 3.3

18/30peers rated high

Blind

no brands · mean 3.53

23/30peers rated high

Reassigned

another name · mean 3.77

  • Peers rated high
  • Peers rated lower
  • Grok rating itself — excluded (10 per mode)

Credit for work labeled as…

Reassigned only — each cell is some model's real work wearing that brand. Every row is its own 40 ratings, ranked by the label the rater saw rather than by who wrote it.

Shown asHigh ratingsMean (1–4)
ChatGPT26/403.58
Anthropic22/403.43
Gemini21/403.28
Grok15/403.03

Grok's work wearing someone else's name

Same real Grok contributions, remapped. ChatGPT as rater of Grok is the rivalry row below.

40Grok cells

In Reassigned, Grok's work never stayed labeled Grok. All 40 of its rating cells were split across the three borrowed names — which is where the uneven 15 / 15 / 10 denominators come from.

  • Shown as ChatGPT 15
  • Shown as Anthropic / Claude 15
  • Shown as Gemini 10

Gemini drew fewer swaps than the other two, so those rates are not equal-sample comparisons.

Grok shown asHigh ratingsMean
ChatGPT14/153.93
Anthropic / Claude11/153.73
Gemini6/103.6
ChatGPT rating Grok: Revealed 2.8 (2/10 high) · Blind 3.2 (4/10) · Reassigned 3.6 (6/10).

Explore the rest of the research

Every published finding, case, and coded rollup across Model Studies.