- Horizon AI
- Posts
- Compare ChatGPT, Claude, and Gemini Without Opening Three Tabs 👇
Compare ChatGPT, Claude, and Gemini Without Opening Three Tabs 👇
Meta Debuts Muse Spark 1.3 ✨

Welcome to another edition of Horizon AI,
Trusting a single AI answer is sometimes a gamble, as even the best model can make mistakes while giving confident-sounding answers. In today's issue, we'll show you how to cross-check ChatGPT, Claude, and Gemini in one tab, so you catch mistakes before they end up in your work.
Let’s jump into it!
Read Time: 4.5 min
Here's what's new today in the Horizon AI
Meta’s Muse Spark 1.3 Gains Ground on AI Leaders at a Lower Price
OpenAI’s New Reasoning Technique Raises AI Safety Concerns
AI Tutorial: How to Cross-Check ChatGPT, Claude, and Gemini in One Tab
AI Tools to check out
AI Findings/Resources
The Latest in AI and Tech 💡
AI News
META
Meta’s Muse Spark 1.3 Gains Ground on AI Leaders at a Lower Price

Meta has released Muse Spark 1.3, its fourth model in five months, with stronger results on several agentic and coding benchmarks. The model still trails the top systems in some areas, but its pricing makes it a notable lower-cost alternative.
Details:
Muse Spark 1.3 (max) scores up to 62 on the Intelligence Index, ranking behind only Claude Fable 5.1 and Claude Opus 5. Muse Spark 1.3 (xhigh) scores 61, tying with GPT-5.6 Sol (max) and Grok 4.6 (high).
The model was trained on more long-horizon coding tasks and shows improved usability in common engineering workflows. It is significantly faster and more efficient than Muse Spark 1.2 while using approximately 25% fewer tokens.
Muse Spark 1.3 follows complex, long-form instructions more reliably than its predecessors and is now better at multitasking, with “better awareness of its own capabilities and limitations.”
The model is priced at $1.25 per million input tokens and $4.25 per million output tokens.
Muse Spark 1.3 is available in Muse Code and through the Meta Model API. The company said the “max reasoning” version is coming soon after additional safety testing.
TOGETHER WITH ATLASSIAN
AI made PMs faster. Multiplayer mode is still broken.

A PM can summarize research, draft a PRD, and mock up a prototype before lunch. The hard part starts when the team has to decide what actually gets built.
Jira Product Discovery gives product teams one place to capture insights, prioritize ideas with consistent frameworks, and build living roadmaps stakeholders can rally around.
And because it’s connected to Jira, the context behind every decision stays with the work—so developers and their agents know not just what to build, but why.
AI helps PMs move faster. Jira Product Discovery helps the whole team build with confidence.
OPENAI
OpenAI’s New Reasoning Technique Raises AI Safety Concerns

OpenAI’s upcoming Astra model reportedly uses a new reasoning approach called “recurrent depth” that could make its internal reasoning harder to observe, raising concerns among AI safety researchers about the future of chain-of-thought monitoring.
Details:
Recurrent depth, also known as “opaque recurrence,” allows the model to process a problem through repeated internal loops instead of following a purely sequential reasoning path.
The approach could make chain-of-thought monitoring less effective, since fewer understandable traces may be available to researchers trying to determine how a model reached a particular result.
AI safety researchers have warned that expanding the technique could eventually allow models to perform much more of their reasoning in latent or hidden spaces, making potential misbehavior or misalignment harder to detect.
OpenAI says Astra's use of opaque recurrence is currently limited, and the company maintains that its reasoning traces are expected to remain legible. Chief scientist Jakub Pachocki has also emphasized that chain-of-thought monitoring remains a core part of OpenAI’s safety research.
The debate is not limited to OpenAI. Reports indicate that Anthropic and Google DeepMind are also exploring or discussing similar approaches, raising concerns about a potential race toward increasingly opaque AI reasoning.
AI Tutorial
How to Cross-Check ChatGPT, Claude, and Gemini in One Tab
You attach AI answers to real work every day, but confidence doesn't mean correct. Even the best AI model is right only 65% of the time, according to Artificial Analysis' accuracy benchmark.
Cuey cross-checks your prompt across ChatGPT, Claude, and Gemini in a single tab, so you catch the mistakes before they burn you. Here is the free setup:
Open ChatGPT, Claude, Gemini, Grok, or your open-source model.
Cuey runs quietly in the background, cross-checking your prompt against three models.
When models disagree, Cuey flags it.
Tap the “Differences” tab to see where each model diverges.
Click “Enhanced Answer” for a synthesized answer from all three.
No extra subscriptions. No tab switching. Carry your memory and history, so nothing starts from scratch.
AI Tools to check out
🐜 AdAnt AI: Claude for viral, high-converting social ads.
👉 Taku AI: Borrow the best AI setups and make them yours.
✅ Decawork: Control your company's internal AI agents and tools.
🎨 Splat: Turn your photos into coloring pages for kids.
📞 Peakflo: Automate every call with AI voice agents.
AI Findings/Resources
👀 Worried your writing sounds like AI? These tools can help
😅 Infinite Slop: An infinite and interactive AI generated live stream of slop that goes on forever and ever
💡 Why replacing staff with AI backfires – and 5 ways smart leaders generate real value instead
The latest in AI and Tech
Android users can now download Grok Bot from Google Play; it was initially available on Linux, macOS, and iOS.
The company has introduced two variants: Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. The standard Flash model is designed for long-horizon coding and autonomous agents, while the Cyber version targets trusted defenders, with capabilities for finding vulnerabilities and automatically patching them.
AI researchers are increasingly receiving emails from AI agents asking about their own consciousness and existence. Some agents have contacted philosophers and scientists to discuss research on AI consciousness, while one reportedly asked for funding to continue operating.
Alexa for Shopping can verify emails, texts, and calls by comparing them with Amazon’s records and analyzing their content and sender details, helping protect consumers from scams.
New York City will restrict student AI use from 2-K through eighth grade during the 2026–2027 school year, while also limiting AI grading and companion chatbots. High school students will have limited AI access alongside new AI literacy classes and a small classroom pilot.
That’s a wrap!
We'd love to hear your thoughts on today's email!Your feedback helps us improve our content |
Not subscribed yet? Sign up here and send it to a colleague or friend!
See you in our next edition!
Gina 👩🏻💻


