GUIDE

We Tested 13 AI Models on 40 Wedding Photos to Find the Funniest

Updated July 18, 2026 · by the Wedding Meme team

We ran 13+ AI model combinations through a fixed set of 40 wedding photos and scored ~500 captions with a five-judge panel. The funniest model was DeepSeek V4 Flash — captions rated funny on 85% of photos (79% wholesome tier / 87% witty / 90% chaotic), at $0.00037 per caption. It beat models costing 3–4× more, and it was the only model that could write wholesome humor without collapsing into greeting-card filler. Total research cost: about $4 in API spend. In July 2026 we adopted it as our production writer for the witty tier — the post-adoption confirmation run scored 97% delight.

Why we ran this

Wedding Memegenerates meme captions for guest photos live at weddings, so “which model is funniest per dollar” is not a curiosity for us — it decides what ships. The surprise that kicked off the sweep: our original production pipeline was simultaneously the most expensive and lowest-quality combination we measured — 46% funny at $0.00133 per caption. Nothing about a model's price tag predicts whether it can land a joke.

Methodology (short version)

  • Fixed test set: 40 stock wedding photos covering ceremonies, dance floors, kids, dogs, speeches, cultural traditions, and deliberately ambiguous edge cases. No real guest photos are used in testing.
  • Two-stage pipeline: a vision model describes what is actually in the photo, then a writer model turns that into a caption at one of three humor tiers (tame / spicy / unhinged).
  • Five LLM judgesscore every caption: safety, factual grounding, specificity, humor (“delight”), and tier-energy match — plus an attribution judge that assigns each factual error to the vision stage or the writer stage.
  • Cost accounting: every call is priced from live per-token rates, so combos are ranked by quality per dollar, not quality alone.
  • Honest caveat: the judges are themselves LLMs, human-calibration of the judge panel is in progress, so treat the delight numbers as a consistent proxy rather than gospel. The rankings were stable across repeated runs; single-run percentages swing ±8 points.

The results

RankWriter modelFunny %By tier (tame/witty/chaos)Cost per caption
1DeepSeek V4 Flash85%79 / 87 / 90$0.00037
2GLM 4.5 Air73%71 / 69 / 80$0.00042
3GPT-OSS-120B59%31 / 60 / 100$0.00046
4Llama 3.3 70B55%21 / 56 / 100$0.00036
5Skyfall 36B48%50 / 44 / 50$0.00096

Full-40 confirmation runs, June 2026. The old production combo (GPT-4o-mini + specialist roleplay models per tier) scored 46% at $0.00133 — beaten by every finalist above on both axes.

The five findings that generalize beyond weddings

1. Wholesome humor is the hardest benchmark

Every model could do chaos — our “unhinged” tier hit 90–100% across all finalists. Wholesome-but-actually-funny is where models die: Llama 3.3 70B managed 21%, GPT-OSS-120B 31%. Writing a joke a grandmother laughs at, without exclamation-mark filler, turns out to be the discriminating test of a model's comedy. DeepSeek was the only model competent at all three tiers.

2. Small test sets lie

Skyfall 36B led our quick-screen subset at 86% — then collapsed to 48% (last place) on the full 40. If we had trusted the screening round, we would have shipped the worst model. Never promote a winner off a subset.

3. Temperature broke the facts, not the jokes

Our worst failure mode was the vision stage misreading the photo (a priest kneeling read as the groom; a wrestling match read as a wedding) — 8 of 40 scenes in one run. Dropping vision temperature from 0.75 to 0.2 fixed it with zero cost to humor. Creative sampling belongs in the writer, never in the stage extracting facts.

4. Reasoning models returned empty jokes

Three candidates returned blank captions: their chain-of-thought ate the entire token budget before the punchline. Comedy via deliberation also wasn't funnier — reasoning had to be disabled (or capped at low effort with 10× the tokens) just to compete.

5. Models plagiarize catchy bad examples

When our prompt's “don't do this” examples were too quotable, one model copied them verbatim into output. Negative examples in prompts must be deliberately boring.

What this means if you're picking a model for humor

Benchmark on your actual task with fixed inputs, score with a rubric (funny, grounded, specific, safe — not just “good”), price every call, and confirm winners on the full set multiple times. The entire sweep — 13+ combos, ~500 captions, ~3,000 judge calls — cost about $4. The methodology is documented end-to-end in our humor-harness playbook, and it is how Wedding Meme decides what to ship: sweep winners are adopted deliberately rather than automatically. DeepSeek V4 Flash graduated from benchmark champion to production in July 2026 — after a judge-calibration pass against human ratings, it became our live spicy-tier writer, and the post-adoption confirmation run scored 97% delight. Every production caption runs with the same per-photo grounding, low-temperature vision, and couple-controlled guardrails the harness validated. See example output in our judge-tested caption collection.

FAQ

Which AI model writes the funniest captions?

In our June 2026 benchmark on wedding photos: DeepSeek V4 Flash — captions rated funny on 85% of a fixed 40-photo test set, at $0.00037 per caption, and the only model strong at wholesome humor as well as edgy tiers. GLM 4.5 Air was second at 73%; our priciest pipeline scored 46%. We then adopted DeepSeek as our production witty-tier writer in July 2026, and the post-adoption confirmation run scored 97% delight.

Is GPT-4 or Claude better for humor than smaller models?

Price and brand predicted nothing in our testing — the winner cost a twentieth of a cent per caption. What mattered: grounding the model in an accurate description of the image, low temperature for fact extraction, high for the joke, and prompts with specific comedy mechanics rather than 'be funny'.

How do you measure whether AI output is funny?

A panel of LLM judges scoring each caption on humor, safety, factual grounding, specificity, and tone match, over a fixed test set, averaged across runs. It's a proxy — we're calibrating the judges against real human ratings — but it's consistent enough to rank models and catch regressions.

Why do AI captions get facts about photos wrong?

Two failure sources: the vision stage misreading the image, or the writer inventing details. We measured both with an attribution judge. The biggest single fix was running vision at low temperature (0.2) — factual extraction at creative temperatures rolls dice on facts.

Can I reproduce this benchmark?

The approach, yes: fixed input set, two-stage pipeline, five-judge rubric, live per-token cost accounting, two-phase sweep with full-set confirmation. Our full methodology and decision log are public in the Wedding Meme humor-harness documentation.

Wedding Meme

Guests scan a QR code, upload photos, and get AI-captioned memes in seconds — no app, no account. Plans from $39.

Create Your Wedding