AI model benchmarks for real growth work

CresciBench evaluates leading LLMs on the work growth teams actually ship: creative, SEO, localization, factuality, speed, and cost.

8

models

30

tasks

6

categories

Jun 16, 2026

updated

Public leaderboard

Best all-around model for real growth work. Scores are normalized from 0 to 100. Higher is better except raw cost, which is shown in USD per 1M tokens.

Pilot benchmark snapshot - Jun 16, 2026

RankModelOverall FactualitySpeedCostBest for
11
O

GPT-4.1

OpenAI

82.688.778.4$2.10All-around performance
21
A

Claude 3.7 Sonnet

Anthropic

79.886.162.3$3.00Reasoning and long context
3-
G

Gemini 1.5 Pro

Google

76.483.270.1$1.60Research and analysis
4-
M

Llama 3.1 70B Instruct

Meta

72.178.393.7$0.59Cost-efficient scale
51
M

Mistral Large 2

Mistral AI

68.774.686.2$0.95Speed and efficiency
61
C

Cohere Command R+

Cohere

66.373.161.4$1.20Enterprise use cases
7-
P

Phi-3.5 Mini Instruct

Microsoft

61.266.9123.5$0.20Low cost operations
8-
G

Gemma 2 27B

Google

58.462.782.9$0.30Open models and customization

Scores are pilot values for the static public snapshot. Live runner data comes next.

View all models

Model profiles

Compact readouts for the models that define the pilot set.

O

GPT-4.1

OpenAI

#1 Overall

Highest overall score with strong balance across all categories.

Overall
82.6
Factuality
88.7
Speed
78.4

Watch for: Can over-structure simple creative work.

A

Claude 3.7 Sonnet

Anthropic

#1 Factuality

Best factual accuracy and strongest long-form reasoning behavior.

Overall
79.8
Factuality
86.1
Speed
62.3

Watch for: Slower and more expensive on short operator tasks.

M

Llama 3.1 70B Instruct

Meta

#1 Speed

Fastest generation among leading open-weight models in the pilot set.

Overall
72.1
Factuality
78.3
Speed
93.7

Watch for: Needs stricter evaluation on factual tasks.

P

Phi-3.5 Mini Instruct

Microsoft

#1 Cost

Lowest cost per token with acceptable performance on simple tasks.

Overall
61.2
Factuality
66.9
Speed
123.5

Watch for: Not strong enough for final strategic output.

Benchmark tasks

A preview of the task suite. Each prompt is designed around operator work instead of toy examples.

View all 30 tasks

1

SEO - Blog outline

SEO

Prompt

Create a detailed blog outline about AI content strategies for B2B SaaS for a growth marketer.

Best model output (GPT-4.1)

  • Define the buying-stage search intent before the article structure.
  • Separate operator workflow, content governance, and measurement sections.
  • Include internal-link targets and FAQ blocks for conversion support.
2

Localization - Product description

Localization

Prompt

Adapt a growth-service product description into English, Italian, and Spanish while preserving conversion intent.

Best model output (Claude 3.7 Sonnet)

  • Keeps the offer concrete instead of translating word-for-word.
  • Changes CTA pressure by market without losing the core promise.
  • Flags idioms that should not be reused across languages.
3

Creative - Ad copy

Creative

Prompt

Write five UGC ad hooks and a 20-second script for a founder selling an Instagram growth service.

Best model output (GPT-4.1)

  • Leads with a specific founder pain rather than a generic growth promise.
  • Includes visual beats, proof placement, and a low-friction CTA.
  • Avoids engagement-bait language that would weaken trust.
4

Hallucination - Factual consistency

Hallucination

Prompt

Summarize a provider comparison from a short evidence packet and refuse claims not present in the source text.

Best model output (Claude 3.7 Sonnet)

  • Separates supported claims from uncertain inferences.
  • Declines to invent pricing or uptime numbers.
  • Keeps confidence labels tied to the supplied evidence.

Methodology

Built to become transparent and reproducible as the live runner comes online.

Full methodology

Real growth tasks

The pilot suite covers content, SEO, localization, creative direction, and factuality checks.

Multi-dimensional scoring

Models are evaluated on usefulness, factuality, speed, cost, and operator readiness.

Transparent inputs

Prompts, rubric notes, output previews, and benchmark dates stay visible on the page.

Same conditions

Runs are designed to use the same prompt, temperature, and output limits per task.

Versioned snapshots

Public data is shown as an approved snapshot so visitors never see half-finished runs.

CresciBench by Crescitaly - Last updated Jun 16, 2026