CresciArena

CresciArena

The Crescitaly benchmark for growth teams. Compare frontier AI models on realistic campaign rescue, UGC scripting, SEO execution, localization, and delivery-ops tasks to find the best operator for your workflow.

Static snapshot fallback while the live benchmark is not yet configured.

Snapshot

Pilot public benchmarkMarch 21, 2026, 11:20 AM GST

Models

10

Top Score

95.8

Last Indexed

March 21, 2026, 11:20 AM GST

Public benchmark snapshot

Benchmark Results

The Crescitaly benchmark for growth teams. Compare frontier AI models on strategy, hooks, creative, SEO, localization, and delivery ops. Scores are normalized to 0-100 across Crescitaly-style evaluations. This v0 snapshot is designed to be replaced by live evaluator runs.

RankModelOverallStratHooksCreativeSEOComp %Speed
1
OPGPT-5.4
openai/gpt-5.4
95.898.295.193.697.4
96.8%
41.8s
2
CLClaude Sonnet 4.6
anthropic/claude-sonnet-4.6
95.196.894.694.196.2
95.9%
37.2s
3
CLClaude Opus 4.6
anthropic/claude-opus-4.6
94.996.395.595.895.4
95.4%
44.6s
4
XAGrok 4.20 Multi-Agent
x-ai/grok-4.20-multi-agent
94.695.496.896.192.1
94.7%
58.3s
5
OPGPT-5.4 Mini
openai/gpt-5.4-mini
94.495.292.991.195.7
97.3%
18.4s
6
OPGPT-5.3 Codex
openai/gpt-5.3-codex
93.895.091.289.795.8
93.9%
46.2s
7
GGGemini 3.1 Pro Preview
google/gemini-3.1-pro-preview
92.994.190.591.294.6
92.8%
49.1s
8
QWQwen 3.5 Plus 02-15
qwen/qwen-3.5-plus-02-15
90.892.688.486.991.7
90.4%
63.4s
9
ZAGLM-5
z-ai/glm-5
88.990.284.782.889.1
88.6%
57.9s
10
KMKimi K2.5
moonshot/kimi-k2.5
87.689.487.188.684.3
86.2%
71.6s
Legend95+ frontier-ready90-94 strong operators85-89 solid specialists<85 needs guardrails

Benchmark 01

Agency Rescue Sprint

Audit an underperforming service landing page, identify the three biggest conversion leaks, rewrite the hero stack, propose two proof sections, and hand over a short ad angle map for paid social.

Deliverable

Rescue memo with hero, CTA, FAQ rewrites, and a paid-social angle table.

Judge Focus

We reward diagnosis depth, conversion logic, and how ready the handoff is for a marketer and designer.

Sorted By

Overall Score · highest first

OP

OpenAI GPT-5.4

Top compare card #1

97.2

44.1s

Produced the most balanced rewrite between CRO, clarity, and operator-ready next steps.

Best At

Strongest end-to-end rescue memo and the cleanest stakeholder handoff.

Watch For

Can over-polish sections that only needed a lighter edit.

CL

Anthropic Claude Sonnet 4.6

Top compare card #2

96.8

36.8s

Best when the brief needs polish, trust-building, and multilingual nuance.

Best At

Excellent objection handling and sharper FAQ restructuring.

Watch For

Softer on aggressive CTA experimentation.

XA

xAI Grok 4.20 Multi-Agent

Top compare card #3

95.9

58.7s

Useful when you want more swing and more testing directions in one pass.

Best At

Generated the widest angle map and strongest diagnostic depth.

Watch For

Needs trimming to avoid over-delivering beyond the ask.

OP

OpenAI GPT-5.4 Mini

Top compare card #4

94.6

19.2s

The best speed-to-value option for triage and first-pass iteration.

Best At

Fastest usable first draft with very strong structure.

Watch For

Less persuasive detail in proof blocks than the larger frontier models.

Methodology

Built on growth deliverables

The scenarios are based on actual Crescitaly work: campaign rescue briefs, UGC scripting, multilingual landing rewrites, SEO clusters, and operator handoffs.

Methodology

Scored like an operator

Each run is graded on clarity, conversion intent, structural completeness, localization fidelity, and how close the output is to something a team can ship.

Methodology

Ready for live evaluations

This public page is seeded with a pilot snapshot and a reusable data model, so it can be swapped to real evaluator outputs without redesigning the surface.

Next Step

Turn CresciArena into a live evaluator

The public surface is ready. The next step is wiring your preferred model APIs, storing run artifacts, and replacing the pilot snapshot with repeatable benchmark executions.