Opus 5.5 Turns One Prompt Into Serious Motion Design

Opus 5.5 shipped this week, and its strongest early reviews focus on a skill that has usually required several tools: making polished explainer videos and motion graphics. On Max effort, the model can create a dynamic 15-second motion graphics video from one prompt, turning a simple instruction into a finished visual sequence.
The model is already the number-one Anthropic model on OpenRouter by share of spend and share of tokens. The switch from Opus 5 has also been particularly rapid, suggesting that users are moving to the new release for more than small quality improvements.
One prompt can now produce a motion graphic video, and the results have drawn enthusiastic reactions. Stephan Livera asked Opus 5.5 to “make a dynamic 15-second motion graphics video that shows what an incredible motion designer you are, like it’s your showreel for a résumé. go all out.”
The response led to reactions from zero, who wrote, “opus 5.5 is f*cking cracked at motion design this entire video is code, 0 after effects im open sourcing the prompt template for these motion designs steal it to recreate these.” Vlad // launch videos posted, “what the fuck… Opus 5.5, extra effort. We are done. This time for real.”
From short clips to complete creative demos
Opus 5.5 is not limited to short visual experiments. It also made a 90-second motion design and sound engineering demo, composing its own piano score for the piece. Klöss described the result as “straight up insane” and asked, “WTF did they feed Opus 5.5?”
Klöss also called the project “a 90-second motion design + sound engineering demo” and said the model “even composed its own piano score.” Taoki responded, “this is actually just straight up good shit. genuine art. what is happening.” Tyler Shukert wrote, “Tried it with Supabase. Amazing results!”
Those reactions point to the main change with Opus 5.5: the model can handle a complete creative task instead of producing only one isolated element. The same system can create motion graphics, assemble the visual sequence, and add a piano score within the demonstrated work.
Opus 5.5 leads several evaluations
Beyond video and motion design, Opus 5.5 now leads SimpleBench at 88.4%. On vision evaluations, it is Anthropic’s best vision model to date. It performs better than Fable 5 and GPT-6 Sol, but worse than GPT-6 Astra, while costing about 60% less than Fable 5.1.
Its performance changes with the effort setting. On Terminal-Bench-Science, Opus 5.5 climbs from 24% at low effort to 62% at xhigh, then drops to 59% at max. GPT-6 Astra and Opus 5.5 lead Fable 5.1 by about 20 points on that benchmark, while Qwen3.8 Max is the best model outside those two labs at 12%.
Many users now say the $200 Claude Code plan beats Codex, while Astra remains the preferred model for review and audit work. Astra also reportedly beat NetHack on its third try and has an 82.5% win rate in DOOM agent matches.
A crowded field keeps moving
Other models are filling different positions in the competition. Luna [Max] entered Code Arena WebDev at number 24, 74 points above GPT-5.6 Luna, at about $0.40 per million tokens blended. Gemini 3.8 Flash scores 41 on the AA Intelligence Index, runs at 291 tokens per second, and supports a 1M context.
Gemini 3.8 Flash posts 89.2% on v2 ARC-AGI at $0.40 per task and 98.5% on v1. On v3, it scores 10.4% with the standard harness and 35% with the provider harness. Xiaomi MiMo-V2.6-Pro is omni-modal with a 1M context, scores 46 on the AA Intelligence Index, and costs $0.13 per task.
Grok 4.7 debuted at number 16 in Agent Arena at $1.14 per task. Meta’s Muse Spark 1.3 is available on GCP and Oracle, while Spark 1.4 has appeared on OpenCode. The range of results shows why model comparisons now depend on the task, the test setup, the effort level, and the price.
Jev targets low-cost decisions
TypeSafe’s Jev is reportedly raising over $1 billion at a valuation above $10 billion. Jev is trained with reinforcement learning for Calibrated Decisions and returns typed decisions with probabilities. It costs $0.044 per 1,000 judgments at 152 milliseconds median latency and is the top model at 1K–10K context on OpenRouter.
Jev stays within 3 points on RewardBench and HaluEval, but trails by 14.5 points on JudgeBench. A cascade that sends low-confidence calls to GPT-6 Astra keeps 99% of accuracy at 57% of the cost, showing how a cheaper model can handle routine decisions while Astra handles uncertain cases.
Ramp matched GPT-5.6 Luna reranking accuracy with 10 times lower tail latency and three times lower cost. Jev also proved 140 Software Foundations theorems for under $1, about 130 times cheaper than Astra.
Opus 5.5 is therefore arriving in a market where raw benchmark scores are only part of the story. Its motion design results, vision performance, coding evaluations, and fast adoption give it a strong opening, while models such as Astra, Gemini 3.8 Flash, Xiaomi MiMo-V2.6-Pro, and Jev continue to compete through different combinations of capability, speed, and cost.




