GPT-4.1
OpenAI
Highest overall score with strong balance across all categories.
- Overall
- 82.6
- Factuality
- 88.7
- Speed
- 78.4
Watch for: Can over-structure simple creative work.
CresciBench evaluates leading LLMs on the work growth teams actually ship: creative, SEO, localization, factuality, speed, and cost.
8
models
30
tasks
6
categories
Jun 16, 2026
updated
Best all-around model for real growth work. Scores are normalized from 0 to 100. Higher is better except raw cost, which is shown in USD per 1M tokens.
Pilot benchmark snapshot - Jun 16, 2026
| Rank | Model | Overall ↓ | Factuality | Speed | Cost | Best for |
|---|---|---|---|---|---|---|
11 | O GPT-4.1 OpenAI | 82.6 | 88.7 | 78.4 | $2.10 | All-around performance |
21 | A Claude 3.7 Sonnet Anthropic | 79.8 | 86.1 | 62.3 | $3.00 | Reasoning and long context |
3- | G Gemini 1.5 Pro | 76.4 | 83.2 | 70.1 | $1.60 | Research and analysis |
4- | M Llama 3.1 70B Instruct Meta | 72.1 | 78.3 | 93.7 | $0.59 | Cost-efficient scale |
51 | M Mistral Large 2 Mistral AI | 68.7 | 74.6 | 86.2 | $0.95 | Speed and efficiency |
61 | C Cohere Command R+ Cohere | 66.3 | 73.1 | 61.4 | $1.20 | Enterprise use cases |
7- | P Phi-3.5 Mini Instruct Microsoft | 61.2 | 66.9 | 123.5 | $0.20 | Low cost operations |
8- | G Gemma 2 27B | 58.4 | 62.7 | 82.9 | $0.30 | Open models and customization |
Scores are pilot values for the static public snapshot. Live runner data comes next.
View all modelsCompact readouts for the models that define the pilot set.
OpenAI
Highest overall score with strong balance across all categories.
Watch for: Can over-structure simple creative work.
Anthropic
Best factual accuracy and strongest long-form reasoning behavior.
Watch for: Slower and more expensive on short operator tasks.
Meta
Fastest generation among leading open-weight models in the pilot set.
Watch for: Needs stricter evaluation on factual tasks.
Microsoft
Lowest cost per token with acceptable performance on simple tasks.
Watch for: Not strong enough for final strategic output.
A preview of the task suite. Each prompt is designed around operator work instead of toy examples.
View all 30 tasks
Prompt
Create a detailed blog outline about AI content strategies for B2B SaaS for a growth marketer.
Best model output (GPT-4.1)
Prompt
Adapt a growth-service product description into English, Italian, and Spanish while preserving conversion intent.
Best model output (Claude 3.7 Sonnet)
Prompt
Write five UGC ad hooks and a 20-second script for a founder selling an Instagram growth service.
Best model output (GPT-4.1)
Prompt
Summarize a provider comparison from a short evidence packet and refuse claims not present in the source text.
Best model output (Claude 3.7 Sonnet)
Built to become transparent and reproducible as the live runner comes online.
The pilot suite covers content, SEO, localization, creative direction, and factuality checks.
Models are evaluated on usefulness, factuality, speed, cost, and operator readiness.
Prompts, rubric notes, output previews, and benchmark dates stay visible on the page.
Runs are designed to use the same prompt, temperature, and output limits per task.
Public data is shown as an approved snapshot so visitors never see half-finished runs.