Discussion about this post

User's avatar
Mark Cleary's avatar

Simon – I'm Art Dir of Short+Sweet since 2002; we run 10 min theatre, film festivals in 14 countries. We run AI agents & built a safety framework since nothing avail that could paste into a system prompt. Couldn't find an email for you. Would really like to know if it holds up against injection patterns you've been tracking. curiositycat online mark@shortandsweet.org

Mira's avatar

The distinction Meta draws between “Instant,” “Thinking,” and a future “Contemplating” mode is interesting, but it also feels like a product workaround for unclear latency and reliability tradeoffs rather than a genuinely new capability split. Their own admission that Muse Spark still has gaps in “long-horizon agentic systems and coding workflows” undercuts the benchmark story more than the Opus/Gemini/GPT comparisons help it. I’m also skeptical that requiring a Facebook or Instagram login for the only public way to test it won’t skew who actually evaluates the model and how seriously developers take it. In practice, access and trust constraints may matter more here than another round of selective benchmark positioning.

4 more comments...

No posts

Ready for more?