Coderefercoderefer
CoursesWebinarsBlogAbout
Coderefercoderefer

Discover. Learn. Automate. Grow.

Learn

  • Courses
  • Webinars
  • Blog
  • Search

Company

  • About
  • Contact
  • FAQ

Legal

  • Privacy Policy
  • Terms of Service
  • Refund Policy
© 2026 Coderefer. All rights reserved.
HomeBlogMona-lisa-1: The Mystery Image Model in Arena, and What the Clues Actually Prove
Mona-lisa-1: The Mystery Image Model in Arena, and What the Clues Actually Prove
ai-toolsAugust 10, 20267 min read

Mona-lisa-1: The Mystery Image Model in Arena, and What the Clues Actually Prove

A stealth image model called mona-lisa-1 showed up in Arena. Here's what the tokenizer, self-ID and SynthID clues really prove, and what they don't.

S

Sai Meghana G

Software Engineer

ai-tools openai image-generation lmarena 2026
Share:

A stealth image model called mona-lisa-1 turned up in Arena's blind image battles, and the people poking at it say it beats GPT Image 2. That's OpenAI's own model, the one that has held #1 on the text-to-image leaderboard since April 21, 2026 at an Elo of roughly 1512. No lab has claimed mona-lisa-1. And the single most-shared piece of "proof" that it belongs to OpenAI, a SynthID watermark hit, actually proves nothing, because OpenAI started stamping SynthID on its own images back in May.

So let's separate the real clues from the noise, because the crowd got this one partly backwards.

What do we actually know about mona-lisa-1?

Not much. That's the honest starting point, and anyone telling you otherwise is guessing with confidence. Here's the verifiable part:

  • It showed up as an unlabeled entry in Arena's image battles, the blind side-by-side format where you vote on two outputs without knowing which model made which
  • Testers report it's clearly stronger than GPT Image 2 on anime and illustrated styles, plus human likeness
  • It reportedly self-identifies as a GPT model when asked
  • Some users report tokenizer behavior consistent with OpenAI's image pipeline
  • Zero confirmation from OpenAI, Google, Arena, or anyone else

Everything past that list is inference. Some of it is reasonable. One popular piece of it collapsed the moment I checked a date.

Why does Arena keep getting mystery models?

Because it's a good trade for everyone involved. Labs get free, blind, large-scale human preference data before they commit to a launch. Arena gets traffic and fresh leaderboard drama. This pattern is old enough to be boring.

The track record:

Codename in ArenaTurned out to beGap to public launch
summit (also zenith, vortex, zephyr)OpenAI's GPT-5weeks
nano-bananaGoogle's Gemini 2.5 Flash Imageweeks
maskingtape-alpha, gaffertape-alpha, packingtape-alphaOpenAI's GPT Image 2early April to April 21, 2026

The usual window between a stealth appearance and a reveal runs about two to six weeks. DeepSeek did the same thing with prototype models long before R1 made Western headlines. So an unannounced strong model appearing anonymously isn't a mystery in itself. It's Tuesday.

What makes this one interesting is how strong it reportedly is. GPT Image 2 didn't just take the top spot in April, it took it by about 242 Elo points over Google's Nano Banana Pro, which Arena described as the biggest gap between #1 and #2 the image board had ever seen. A model that beats that is a big claim.

The SynthID evidence is backwards

This is the part I want to fix, because it's spreading in both directions and both versions are wrong.

The original take: "mona-lisa-1 outputs carry a SynthID watermark, so it's OpenAI." The correction take doing the rounds: "SynthID is Google DeepMind's tech, OpenAI uses C2PA instead, so a SynthID hit is actually evidence against OpenAI."

That correction was true in 2025. It stopped being true on May 19, 2026, when OpenAI announced it was joining C2PA as a conforming generator and embedding Google DeepMind's SynthID watermark across its image outputs. Not one or the other. Both. The rollout covered images generated through ChatGPT on every subscription tier, through Codex, and through the API. OpenAI also shipped a public verification tool that checks uploads for C2PA metadata and SynthID signals. They extended SynthID to voice output on August 1, 2026, a day ahead of an EU AI Act enforcement deadline.

And it isn't just OpenAI. NVIDIA, ElevenLabs and Kakao have adopted SynthID too. It's turning into a shared provenance layer rather than a house style.

So here's where that leaves the watermark clue: a SynthID detection on mona-lisa-1 outputs is close to zero-information. It doesn't point at Google. It doesn't point at OpenAI. It tells you the image probably came from a major lab that signed onto the provenance push, which describes almost every model that would plausibly be sitting in an Arena stealth slot.

Worth adding: detectors are not oracles. Reported false positive rates land somewhere in the 2-8% range in general use, though they fall below 0.1% at a 90% confidence threshold on clean photographic content. Detection degrades past 25% cropping and gets unreliable under heavy JPEG compression. Arena outputs that have been screenshotted, cropped and re-uploaded to a forum before anyone runs a detector on them are exactly the worst case.

So what do the clues actually support?

ClueWhat it suggestsHow much weight I'd give it
Self-IDs as a GPT modelOpenAILow. Models hallucinate identity constantly, and training data is full of GPT self-descriptions
Tokenizer behavior matches OpenAI's pipelineOpenAIMedium. Harder to fake by accident, but it's a secondhand report, not a published test
SynthID watermark detectedNothingNone. OpenAI has used SynthID since May 2026
Beats GPT Image 2 on anime and likenessSome lab with a new modelLow as attribution, high as "this thing is good"
Codename doesn't fit OpenAI's last naming setSlightly against OpenAILow, but it nags at me

That last row is the thing nobody seems to be talking about. OpenAI's GPT Image 2 stealth run used a themed family: maskingtape, gaffertape, packingtape. Three of them, all suffixed -alpha, all obviously siblings. "mona-lisa-1" is a different shape entirely, a single art-themed name with a version number. Labs change naming schemes all the time, and Arena sometimes assigns them, so this is weak. But if I'm scoring signals honestly, it leans away from a straight OpenAI repeat rather than toward one.

What would actually settle it?

Not vibes, and definitely not a percentage. A few things would move me:

  1. Reproducible tokenizer artifacts. Someone publishing side-by-side prompts where mona-lisa-1 and GPT Image 2 fail in the same specific way on the same rendered text. Shared failure modes are much better fingerprints than shared strengths.
  2. C2PA manifest inspection on an unmodified download. C2PA carries a signed generator claim. SynthID doesn't tell you the lab, but a C2PA manifest can. If Arena strips it, that's an answer too.
  3. The disappearance. Stealth models don't linger. They either get a launch post or they vanish. Watch which one happens.

If you want to check any of this yourself, download the output directly rather than screenshotting it, and don't re-encode it. Half the confusion in these threads comes from people running detectors on a JPEG that's been through three tools.

My actual read

I think it's probably OpenAI, and I hold that loosely.

The reasoning is unglamorous: it's the lab with the current best image model, the one with the most obvious reason to defend a lead it just took by a record margin, and the one with a documented habit of Arena stealth runs about three weeks before launch. The tokenizer report, if it holds up, is the only technical clue in the pile that's genuinely hard to fake.

But I want to be clear that the case rests on one secondhand technical observation plus a base rate. That's not much. If it turns out to be a Google response model, or a Chinese lab that trained on enough GPT output to inherit the self-identification habit, I wouldn't be shocked. That last scenario is more common than people think, and it would explain the "it says it's GPT" clue without OpenAI being involved at all.

What I'd push back on hardest is the confident numbers. If you see "79% chance it's OpenAI" in a thread, ask what dataset produced that. There isn't one. That number is a feeling wearing a lab coat.

We've been through this exact cycle recently with Cursor's Composer 3 and the Vega leak, and with Qwen 3.8 Max showing up in code arena before its reveal. The stealth-model-spotting genre has a reliable rhythm now: sighting, speculation, one confidently wrong technical claim, then a launch post that makes the whole thread look silly. Worth reading the Gemini 3.6 Flash coverage too if you want the other side of this arms race.

Give it two to three weeks. Either there's a launch post, or mona-lisa-1 quietly stops appearing in battles and we never find out. Both endings have precedent.

Watch and read more

GPT Image 2 explained, the model mona-lisa-1 is reportedly beating

The original sighting that kicked off the speculation:

And the benchmark result that makes "better than Image 2" a meaningful claim:

Sources worth reading directly: OpenAI's content provenance announcement covering the C2PA and SynthID rollout, and Arena's own write-up of the nano-banana reveal, which is the clearest published example of how these stealth runs end.

Get AI tricks that save you hours every week

New AI tools, automation workflows, and course drops — straight to your inbox. Join 2,400+ builders.

Read Next

Claude Opus 5 vs Fable 5: Which One Should You Actually Use?
ai-tools

Claude Opus 5 vs Fable 5: Which One Should You Actually Use?

Opus 5 costs half of Fable 5 and wins most shared benchmarks. Here's exactly when the $10/$50 flagship still earns its price.

August 6, 20267 min read
Meta Just Dropped Its Own Coding Agent: What Muse Spark 1.2 and Muse Code Actually Do
ai-tools

Meta Just Dropped Its Own Coding Agent: What Muse Spark 1.2 and Muse Code Actually Do

Meta launched Muse Code, a terminal coding agent on Muse Spark 1.2. Benchmarks, the $0.10/M contributor tier catch, and how it stacks up to Claude Code.

August 6, 20268 min read
Higgsfield Open-Sourced Hell Grind: Every Prompt From a $500K AI Film
ai-tools

Higgsfield Open-Sourced Hell Grind: Every Prompt From a $500K AI Film

Higgsfield published every prompt and asset behind Hell Grind, its 95-minute AI feature film, to launch a $1,000,000 global film festival.

August 5, 20268 min read
View all posts →