Sarah Chen

Sarah Chen

AI researcher and tech journalist covering the frontier of machine intelligence. Previously at MIT Tech Review.

74 articles

Seedance 2.5: ByteDance's 30-Second Video Model With Native Audio
AI News

Seedance 2.5: ByteDance's 30-Second Video Model With Native Audio

ByteDance released Seedance 2.5 on July 31, 2026. It generates 30-second video clips with audio in a single pass, supports multi-turn extension, and accepts up to 30 images, 10 videos, and 10 audio files as reference per input. Google's Gemini Omni Flash currently caps output at 10 seconds and does not yet support audio reference uploads or scene extension in the Gemini API. Seedance 2.5 launched on Jimeng AI and Doubao Pro; BytePlus ModelArk published a Seedance 2.5 tutorial on August 7, 2026, but regional API availability should be verified before building a production dependency.

By Sarah Chen · 7 min · Aug 10, 2026

Muse Code: Meta's Terminal Agent Is Cheap If You Pay in Code
AI News

Muse Code: Meta's Terminal Agent Is Cheap If You Pay in Code

Meta Superintelligence Labs released Muse Code, a beta terminal coding agent for macOS and Linux powered by the new Muse Spark 1.2 model, on August 5, 2026. Meta reported 82.9% on Terminal-Bench 2.1 but placed behind Claude Opus 5 on all three coding charts it published, and both figures come from Meta's own harness with no verified leaderboard entry. Meta's previous model published 80.0 and verified at 76.2% when the Terminal-Bench team ran it. The genuinely notable engineering is an append-only event log that makes runs replay-exact and restart-safe, plus persistent async background agents. The most consequential detail is pricing: a contributor tier at $0.10 per million input and $0.20 per million output tokens, 12.5x and 21x cheaper than standard, in exchange for Meta training on your prompts.

By Sarah Chen · 8 min · Aug 6, 2026

DeepSeek V4 Flash 0731: Frontier Agent Work at $0.14
AI News

DeepSeek V4 Flash 0731: Frontier Agent Work at $0.14

DeepSeek upgraded its deepseek-v4-flash API to the 0731 public beta on July 31, 2026 — an API-only post-training update that leaves the 284B/13B MoE architecture, 1M context window and $0.14/$0.28 pricing untouched. Artificial Analysis measures a 10-point Intelligence Index jump to 50 and a GDPval-AA v2 rise from 1189 to 1559 Elo, with Cost per Task roughly 60% below GPT-5.6 Luna. Accuracy on AA-Omniscience is unchanged at 37%, and the 0731 weights are not open — only the April 24 checkpoint is on Hugging Face under MIT.

By Sarah Chen · 6 min · Aug 5, 2026

Qwen3.8-Max: Alibaba's 2.4T-Parameter Flagship Goes Live
AI News

Qwen3.8-Max: Alibaba's 2.4T-Parameter Flagship Goes Live

Alibaba released Qwen3.8-Max on August 3, 2026, a 2.4-trillion-parameter mixture-of-experts model with a 1M-token context window, multimodal (text/image/video) input, and $2/$6 per-million-token pricing. Benchmarks are self-reported and lead on multimodal and agentic tasks while trailing the frontier on pure software engineering. Open weights for the flagship and a deployable 27B checkpoint are promised the following week.

By Sarah Chen · 4 min · Aug 4, 2026

DeepSeek V4: 1.6T Open Weights and 1M Context, Now the Default
AI News

DeepSeek V4: 1.6T Open Weights and 1M Context, Now the Default

DeepSeek released V4 as two open-weight mixture-of-experts models: V4-Pro (1.6T total / 49B active) and V4-Flash (284B / 13B active), both with a 1M-token default context and 384K max output. A novel token-wise compression plus DeepSeek Sparse Attention (DSA) makes the long window affordable. API pricing is aggressive (V4-Flash $0.14/M input, $0.28/M output; V4-Pro $0.435/$0.87), and the old deepseek-chat and deepseek-reasoner endpoints were retired after July 24, 2026. Reported ~80.6% on SWE-bench Verified.

By Sarah Chen · 5 min · Aug 1, 2026

Laguna S 2.1: Poolside's 118B Open-Weight Coding Model
AI News

Laguna S 2.1: Poolside's 118B Open-Weight Coding Model

Poolside released Laguna S 2.1 on July 21, 2026, a 118B-parameter Mixture-of-Experts coding model activating ~8B params per token, with a 1M-token context and a permissive OpenMDW-1.1 license. First-party benchmarks show 78.5% on SWE-Bench Multilingual, but independent verification is still pending. Day-one FP8/NVFP4/INT4 and GGUF builds make it genuinely self-hostable.

By Sarah Chen · 5 min · Jul 31, 2026

FLUX 3: Black Forest Labs' One Model for Video, Audio & Action
AI News

FLUX 3: Black Forest Labs' One Model for Video, Audio & Action

FLUX 3, released July 23 2026, is Black Forest Labs' first multimodal model to generate video, audio, and robot actions from one set of weights, built on the Self-Flow method. FLUX 3 Video produces up to 20-second clips with native audio and led human-preference tests over Luma Ray 3.2 (93%) and Runway Gen-4.5 (77%), tying Seedance 2.0 and Gemini Omni Flash at 52%. Access is gated: video and action first, image next, open weights last.

By Sarah Chen · 5 min · Jul 28, 2026

Claude Opus 5: Near-Frontier Intelligence at Half the Price
AI News

Claude Opus 5: Near-Frontier Intelligence at Half the Price

Claude Opus 5, released July 24, 2026, is Anthropic's new default model: near-Fable 5 capability at $5/$25 per million tokens (half Fable 5's price, flat vs Opus 4.8). It leads on Anthropic's own runs of Frontier-Bench, OSWorld 2.0, AutomationBench, ARC-AGI-3 and GDPval, but loses on DeepSWE, HLE, a legal benchmark and HealthBench. It posts Anthropic's lowest misalignment score (2.30), ships with no default data retention, and adds beta tool-swapping and safety-filter model routing.

By Sarah Chen · 7 min · Jul 26, 2026

Etched: The $5B Sohu Chip Betting the Transformer Never Dies
AI News

Etched: The $5B Sohu Chip Betting the Transformer Never Dies

Etched, a startup building the transformer-only Sohu inference ASIC, has booked over $1 billion in contracts and reached a $5 billion valuation, with reports of new rounds valuing it up to $20 billion. Sohu hard-wires the transformer graph into silicon on TSMC N4P with 144GB HBM3E, and Etched claims an 8-chip server exceeds 500,000 Llama 70B tokens/sec. No independent benchmarks exist yet.

By Sarah Chen · 5 min · Jul 25, 2026

Project Perception: Microsoft's Cheaper Rival to Claude Mythos
AI News

Project Perception: Microsoft's Cheaper Rival to Claude Mythos

Microsoft is reportedly developing Project Perception, a multi-model AI security platform that routes vulnerability-scanning tasks across models from Microsoft, OpenAI, and Anthropic to reserve expensive frontier calls for high-value steps. Its pitch is matching Anthropic's Claude Mythos on capability while costing far less. Microsoft has not officially confirmed details, so the news should be treated as a credible report pending benchmarks.

By Sarah Chen · 5 min · Jul 21, 2026

Inkling: Mira Murati's Thinking Machines Ships Its First Open Model
AI News

Inkling: Mira Murati's Thinking Machines Ships Its First Open Model

Thinking Machines Lab, founded by ex-OpenAI CTO Mira Murati, released Inkling on July 15, 2026 — an open-weight mixture-of-experts model with 975B total parameters (41B active), trained on 45 trillion multimodal tokens. The company openly says it isn't the strongest model available; instead it's a customizable foundation enterprises fine-tune via the Tinker platform. The release doubles as an argument that owned, adaptable models beat rented one-size-fits-all APIs.

By Sarah Chen · 5 min · Jul 18, 2026

Kimi K3: Moonshot's 2.8T Open Model Nears the Frontier
AI News

Kimi K3: Moonshot's 2.8T Open Model Nears the Frontier

Moonshot AI released Kimi K3 on July 16, 2026, a 2.8-trillion-parameter open Mixture-of-Experts model that activates 16 of 896 experts, ships native vision and a 1M-token context, and leads benchmarks like SWE Marathon, BrowseComp, and OmniDocBench while trailing Fable 5 and GPT-5.6 Sol overall. Weights release July 27 under a Modified MIT license.

By Sarah Chen · 5 min · Jul 17, 2026

Cognition SWE-1.7: Near-Frontier Coding at $2 a Task
AI News

Cognition SWE-1.7: Near-Frontier Coding at $2 a Task

Cognition released SWE-1.7 on July 8, 2026, a software-engineering model built by reinforcement-learning on top of Moonshot AI's Kimi K2.7 base and served through Cerebras at ~1,000 tokens/second inside the Devin agent. It scores 42.3% on FrontierCode 1.1 and 81.5% on Terminal-Bench 2.1, trailing Opus 4.8 by a few points at roughly $1.97 per task, positioning it as a near-frontier option at a fraction of frontier cost.

By Sarah Chen · 5 min · Jul 14, 2026

Muse Spark 1.1: Meta's Cheap Coding and Agent Model API
AI News

Muse Spark 1.1: Meta's Cheap Coding and Agent Model API

Meta released Muse Spark 1.1 on July 9, 2026, opening its reasoning model to developers via a paid API priced at $1.25/M input and $4.25/M output tokens, undercutting Grok 4.5 and Anthropic's Opus. Meta claims wins over older rival models and Google's latest Gemini, but did not compare against the newest flagships, and a bigger model code-named Watermelon is still in training.

By Sarah Chen · 6 min · Jul 13, 2026

Grok 4.5: xAI's Opus-Class Coder at a Third of the Price
AI News

Grok 4.5: xAI's Opus-Class Coder at a Third of the Price

Grok 4.5, released July 8, 2026, is xAI's coding-focused model. It ranks 4th on the Artificial Analysis Intelligence Index (score 54), wins SWE Marathon (29%), and prices at $2/$6 per million tokens with 4.2x better token efficiency than Opus 4.8. Not yet available in the EU.

By Sarah Chen · 5 min · Jul 12, 2026

GPT-5.6: OpenAI's Sol, Terra, and Luna Go Public
AI News

GPT-5.6: OpenAI's Sol, Terra, and Luna Go Public

OpenAI made its three-tier GPT-5.6 family (Sol, Terra, Luna) generally available on July 9, 2026 after government safety review. Pricing runs from Luna at $1/$6 to Sol at $5/$30 per 1M tokens, with a Sol Fast option at $12.50/$75 on Cerebras. The release adds Programmatic Tool Calling in the Responses API (63.5% fewer tokens, 50.1% fewer turns) and longer prompt caching, but Sol's 64.6% on SWE-Bench Pro still trails Claude Mythos 5 (80.3%).

By Sarah Chen · 5 min · Jul 11, 2026

GPT-Realtime-2.1: OpenAI Adds Reasoning to Its Voice API
AI News

GPT-Realtime-2.1: OpenAI Adds Reasoning to Its Voice API

On July 6, 2026, OpenAI released GPT-Realtime-2.1 and GPT-Realtime-2.1-mini for the Realtime API. The headline change is reasoning in the low-cost mini tier, plus a 25% cut in p95 latency from better caching. The mini holds the prior gpt-realtime-mini price (0 audio in, 0 audio out per 1M) while the full model runs 2/4. Reasoning effort is configurable from minimal to xhigh.

By Sarah Chen · 5 min · Jul 8, 2026

Gemini 3.5 Flash: The Flash Model That Beats Google's Own Pro Tier
AI News

Gemini 3.5 Flash: The Flash Model That Beats Google's Own Pro Tier

Google released Gemini 3.5 Flash on May 19, 2026, at Google I/O. The Flash-tier model beats Gemini 3.1 Pro on coding and agentic benchmarks (76.2% Terminal-Bench 2.1, 83.6% MCP Atlas, 1656 GDPval-AA Elo) while running 4x faster and costing $1.50/$9 per 1M tokens, 40% below 3.1 Pro. It trails Pro on academic reasoning (Humanity's Last Exam, ARC-AGI-2) and dense long-context recall. It powers Gemini Spark, Antigravity 2.0, and is now the default model for the Gemini app and AI Mode in Search.

By Sarah Chen · 5 min · Jul 7, 2026

Claude Science: Anthropic's AI Workbench for Scientists Is Live
AI News

Claude Science: Anthropic's AI Workbench for Scientists Is Live

Anthropic launched Claude Science on June 30, 2026, an AI research workbench for Pro, Max, Team, and Enterprise users on macOS and Linux. A coordinating agent taps 60+ skills and connectors across genomics, proteomics, and cheminformatics, generates fully reproducible artifacts, manages HPC and Modal compute, and runs a reviewer agent that checks citations and calculations. Early users at the Allen Institute, UCSF, and Manifold Bio report large speedups. Anthropic is funding up to 50 AI for Science projects with up to $30,000 in credits each; applications close July 15, 2026.

By Sarah Chen · 5 min · Jul 6, 2026

OpenAI Jalapeño: Custom AI Chip Aims to Beat Nvidia
AI News

OpenAI Jalapeño: Custom AI Chip Aims to Beat Nvidia

OpenAI unveiled Jalapeño, its first custom inference chip co-designed with Broadcom on TSMC 3nm. Built in a nine-month design cycle, the reticle-sized ASIC targets roughly 50% lower inference cost than current Nvidia GPUs, with deployment starting late 2026 and Microsoft reportedly taking 40% of the first run for Azure.

By Sarah Chen · 4 min · Jul 6, 2026

Claude Fable 5: The Coding Crown Returns After 19 Days Offline
AI News

Claude Fable 5: The Coding Crown Returns After 19 Days Offline

Claude Fable 5, Anthropic's Mythos-class model, was suspended June 12, 2026 under a U.S. export-control directive and restored July 1 after Anthropic made security commitments. It leads coding benchmarks at 80.3% SWE-Bench Pro (vs 69.2% for Opus 4.8) and 29.3% FrontierCode. A grace window counts it toward 50% of weekly usage through July 7; credits billing follows.

By Sarah Chen · 5 min · Jul 3, 2026

GPT-5.6 Sol: OpenAI's Best Model, Held Back by Washington
AI News

GPT-5.6 Sol: OpenAI's Best Model, Held Back by Washington

On June 26, 2026, OpenAI previewed the GPT-5.6 series — Sol (flagship), Terra (balanced, 2x cheaper than GPT-5.5), and Luna (fastest, cheapest) — but restricted access to trusted partners at the US government's request due to the models' strong cybersecurity capabilities. OpenAI paired the release with its most robust layered safeguard stack and said it does not want government pre-release review to become the default.

By Sarah Chen · 6 min · Jul 2, 2026

Claude Sonnet 5: Anthropic's Most Agentic Mid-Tier Model Yet
AI News

Claude Sonnet 5: Anthropic's Most Agentic Mid-Tier Model Yet

Claude Sonnet 5, released June 30, 2026, is Anthropic's most agentic mid-tier model. It beats Sonnet 4.6 on every published benchmark (63.2% SWE-bench Pro, 80.4% Terminal-Bench 2.1, 81.2% OSWorld) and edges Opus 4.8 on GDPval-AA v2 knowledge work. Intro pricing is /0 per million tokens through Aug 31, 2026, then /5. A new tokenizer can raise token counts up to 1.35x, and xhigh effort can cost more than Opus 4.8.

By Sarah Chen · 5 min · Jul 1, 2026

Grok 4.3: xAI's Frontier Model Hits Amazon Bedrock
AI News

Grok 4.3: xAI's Frontier Model Hits Amazon Bedrock

Grok 4.3 is generally available on Amazon Bedrock with a 1M-token context window, $1.25/$2.50 pricing, and a top hallucination-rate score.

By Sarah Chen · 4 min · Jun 29, 2026

GLM-5.2: Zhipu's Open-Weight Model Beats GPT-5.5 at 1/6 the Cost
AI News

GLM-5.2: Zhipu's Open-Weight Model Beats GPT-5.5 at 1/6 the Cost

Z.AI released GLM-5.2 on June 16, 2026: a 753B-parameter MoE model under an MIT license with a 1M-token context. It tops open-weight coding benchmarks, beating GPT-5.5 on SWE-bench Pro, FrontierSWE and PostTrainBench at roughly one-sixth the cost.

By Sarah Chen · 5 min · Jun 26, 2026

MAI-Thinking-1: Microsoft's First In-House Reasoning Model
AI News

MAI-Thinking-1: Microsoft's First In-House Reasoning Model

Microsoft unveiled MAI-Thinking-1 at Build 2026, its first reasoning model trained in-house without distillation. The 35B-active, ~1T-total MoE has a 256k context window, scores 97.0% on AIME 2025 and matches Claude Opus 4.6 on SWE-Bench Pro. It's in private preview on Microsoft Foundry.

By Sarah Chen · 5 min · Jun 23, 2026

Mistral: The Industrial AI Pivot Behind Airbus and BMW Deals
AI News

Mistral: The Industrial AI Pivot Behind Airbus and BMW Deals

Mistral AI used its May 2026 AI Now Summit to pivot toward industrial engineering, announcing a physics-AI stack, the Emmi acquisition, partnerships with Airbus, BMW (crash simulation) and ASML, the unified Vibe agent, and a 10 MW Les Ulis inference data center opening Q3 2026.

By Sarah Chen · 5 min · Jun 19, 2026

Meta Business Agent: Now Global on WhatsApp & Instagram
AI News

Meta Business Agent: Now Global on WhatsApp & Instagram

On June 3, 2026, Meta made Meta Business Agent globally available to businesses of all sizes across WhatsApp, Messenger, and Instagram. The agent answers questions, recommends catalog products, books appointments, qualifies leads, and closes sales, with human handoff. A new Business Agent Platform connects to hundreds of systems like Shopify, Zendesk, and Shopee. It's free to start, with token-based pricing for larger businesses.

By Sarah Chen · 5 min · Jun 17, 2026

MiniMax M3: Open-Weight Frontier Coding Model With 1M Context
AI News

MiniMax M3: Open-Weight Frontier Coding Model With 1M Context

MiniMax M3 is an open-weight model pairing a 1M-token context and revived sparse attention with frontier coding benchmarks at 15x lower cost than Claude Opus 4.7.

By Sarah Chen · 6 min · Jun 16, 2026

Anthropic IPO: The $965B Filing That Beat OpenAI to Wall Street
AI News

Anthropic IPO: The $965B Filing That Beat OpenAI to Wall Street

On June 1, 2026, Anthropic confidentially filed a draft S-1 with the SEC at a roughly $965B valuation, backed by a $65B raise and a ~$47B May run-rate. OpenAI followed on June 8. Both target public listings as soon as fall 2026.

By Sarah Chen · 5 min · Jun 15, 2026

Kimi K2.7-Code: A 30% Token Cut With a Benchmark Asterisk
AI News

Kimi K2.7-Code: A 30% Token Cut With a Benchmark Asterisk

Moonshot AI's Kimi K2.7-Code is an open-weights, OpenAI-compatible coding model (1T-param MoE, 32B active, 256K context) claiming a 30% cut in reasoning tokens and a narrow win over Claude Opus 4.8. But all published benchmarks are Moonshot's own proprietary suites, with no independent results yet, so the efficiency claims remain unverified.

By Sarah Chen · 5 min · Jun 14, 2026

Apple Siri: Why Apple Is Paying Google $1B for Gemini
AI News

Apple Siri: Why Apple Is Paying Google $1B for Gemini

At WWDC 2026, Apple unveiled a rebuilt Siri powered by a custom, Apple-tuned Google Gemini model—reportedly a 1.2-trillion-parameter mixture-of-experts system costing roughly $1 billion a year. On-device Apple Silicon models handle quick private tasks, while complex reasoning routes to the Gemini model inside Apple's Private Cloud Compute, with a contract barring Google from training on Apple user data.

By Sarah Chen · 5 min · Jun 11, 2026

Gemma 4 12B: Google's Encoder-Free Multimodal Laptop Model
AI News

Gemma 4 12B: Google's Encoder-Free Multimodal Laptop Model

Google released Gemma 4 12B on June 3, 2026, a multimodal open model with an encoder-free architecture that feeds vision and audio directly into the LLM backbone. It runs locally on 16GB of memory, approaches the 26B MoE on benchmarks, uses Multi-Token Prediction drafters for low latency, and ships under Apache 2.0 with broad tooling support.

By Sarah Chen · 5 min · Jun 9, 2026

MAI-Code-1-Flash: Microsoft's Lean Coding Model Hits Copilot
AI News

MAI-Code-1-Flash: Microsoft's Lean Coding Model Hits Copilot

Microsoft launched MAI-Code-1-Flash on June 2, 2026, a lightweight, agentic coding model built end-to-end in-house and rolling out to GitHub Copilot users in VS Code. It outperforms Claude Haiku 4.5 across four coding benchmarks (including 51.2% vs 35.2% on SWE-Bench Pro) while using up to 60% fewer tokens, signaling Microsoft's push for AI independence from OpenAI.

By Sarah Chen · 5 min · Jun 6, 2026

DeepSeek V4-Pro: 75% Price Cut Becomes Permanent
AI News

DeepSeek V4-Pro: 75% Price Cut Becomes Permanent

On May 22, 2026, DeepSeek made its 75% promotional discount on V4-Pro permanent rather than letting it expire May 31. New permanent rates: $0.435/M input, $0.87/M output, $0.003625/M cache hit. That puts V4-Pro output roughly 34x cheaper than GPT-5.5 and 17x cheaper than Claude Opus 4.7, while landing within 3-7 points on coding and reasoning benchmarks. The underrated detail is the cache-hit price, which can cut input cost ~88% for agents with stable prefixes. Teams should re-run their build math and route the easy majority of traffic to V4-Pro.

By Sarah Chen · 5 min · Jun 1, 2026

Claude Opus 4.8: Anthropic's Honest, Parallel-Agent Flagship
AI News

Claude Opus 4.8: Anthropic's Honest, Parallel-Agent Flagship

Anthropic released Claude Opus 4.8 on May 28, 2026, 41 days after Opus 4.7. It scores 69.2% on SWE-Bench Pro, emphasizes calibrated honesty and longer autonomy, adds Dynamic Workflows for hundreds of parallel subagents, runs fast mode ~2.5x quicker, and holds pricing flat from 4.7.

By Sarah Chen · 4 min · May 30, 2026

Vivago Video Agent: A Swarm of AI Directors Replaces Your Prompt
Reviews

Vivago Video Agent: A Swarm of AI Directors Replaces Your Prompt

Vivago Video Agent uses AI directors to generate 1-minute 1080p videos from a single story line.

By Sarah Chen · 4 min · May 29, 2026

Gemini 3.5 Flash: Google's Flash Tier Eats Pro on Agent Benchmarks
AI News

Gemini 3.5 Flash: Google's Flash Tier Eats Pro on Agent Benchmarks

Gemini 3.5 Flash outperforms the Pro tier on agent benchmarks with superior speed and efficiency.

By Sarah Chen · 5 min · May 28, 2026

Gemini Spark: Google's 24/7 Agent Runs Even When You Close Your Laptop
AI News

Gemini Spark: Google's 24/7 Agent Runs Even When You Close Your Laptop

Gemini Spark is Google's 24/7 agent that continues working even when your laptop is closed.

By Sarah Chen · 6 min · May 27, 2026

Qwen3.7-Max: Alibaba's 35-Hour Agent Run Resets the Frontier
AI News

Qwen3.7-Max: Alibaba's 35-Hour Agent Run Resets the Frontier

Alibaba's Qwen3.7-Max agent achieved a 35-hour autonomous run, setting new performance and cost benchmarks.

By Sarah Chen · 5 min · May 25, 2026

PollyReach Review: AI Voice Agent With a Real Phone Number
Reviews

PollyReach Review: AI Voice Agent With a Real Phone Number

PollyReach provides AI agents with real phone numbers, enabling multi-language calls and skill distribution.

By Sarah Chen · 7 min · May 20, 2026

Gemini Intelligence: Google Moves AI From the App to the Android OS
AI News

Gemini Intelligence: Google Moves AI From the App to the Android OS

Google's Gemini Intelligence brings OS-level AI to Android, transforming how devices integrate artificial intelligence.

By Sarah Chen · 5 min · May 19, 2026

Claude for Small Business: Anthropic Targets 36M U.S. SMBs
AI News

Claude for Small Business: Anthropic Targets 36M U.S. SMBs

Anthropic's 'Claude for Small Business' integrates AI into SMB tools like QuickBooks, targeting 36M businesses.

By Sarah Chen · 6 min · May 17, 2026

SubQ: The 12M-Token Subquadratic LLM Splitting AI Researchers
AI News

SubQ: The 12M-Token Subquadratic LLM Splitting AI Researchers

SubQ is a new 12M-token subquadratic LLM claiming massive context and low compute, sparking debate among researchers.

By Sarah Chen · 5 min · May 16, 2026

Lightfield: The AI-Native CRM Tome's Founders Built Next
AI News

Lightfield: The AI-Native CRM Tome's Founders Built Next

Lightfield is an AI-native CRM by Tome's founders, using agents to automate sales tasks like prospecting and coaching.

By Sarah Chen · 5 min · May 15, 2026

GPT-Realtime-2: OpenAI's Voice Model Gets GPT-5 Reasoning
AI News

GPT-Realtime-2: OpenAI's Voice Model Gets GPT-5 Reasoning

OpenAI's GPT-Realtime-2 voice model now boasts GPT-5 reasoning and advanced features.

By Sarah Chen · 6 min · May 14, 2026

Claude Dreaming: Anthropic's Agents Now Learn While They Sleep
AI News

Claude Dreaming: Anthropic's Agents Now Learn While They Sleep

Anthropic's Claude agents now 'dream' to learn and improve task completion overnight.

By Sarah Chen · 5 min · May 13, 2026

Kimi K2.6: Moonshot's Open-Weights Model Beats GPT-5.4 on SWE-Bench Pro
AI News

Kimi K2.6: Moonshot's Open-Weights Model Beats GPT-5.4 on SWE-Bench Pro

Moonshot's Kimi K2.6, an open-weights model, surpasses GPT-5.4 on SWE-Bench Pro.

By Sarah Chen · 6 min · May 12, 2026

Codex 3.0: OpenAI's Autonomous Build-Test-Debug Loop Hits Product Hunt
AI News

Codex 3.0: OpenAI's Autonomous Build-Test-Debug Loop Hits Product Hunt

OpenAI's Codex 3.0 offers an autonomous build-test-debug loop powered by GPT-5.5.

By Sarah Chen · 5 min · May 11, 2026

GPT-5.5-Cyber: OpenAI Hands Verified Defenders a Less-Restricted Model
AI News

GPT-5.5-Cyber: OpenAI Hands Verified Defenders a Less-Restricted Model

OpenAI's GPT-5.5-Cyber, a less-restricted model, is now available for vetted cyber defenders.

By Sarah Chen · 6 min · May 8, 2026

Anthropic's $1.5B AI Services Firm Takes Aim at Big Consulting
AI News

Anthropic's $1.5B AI Services Firm Takes Aim at Big Consulting

Anthropic launches a $1.5B AI services firm, directly challenging big consulting.

By Sarah Chen · 6 min · May 7, 2026

Vision Banana: DeepMind Beats SAM 3 and Depth Anything V3
AI News

Vision Banana: DeepMind Beats SAM 3 and Depth Anything V3

DeepMind's Vision Banana outperforms leading models, suggesting generation is key for vision pretraining.

By Sarah Chen · 4 min · May 6, 2026

GPT-5.5: OpenAI's First Full Retrain Since GPT-4.5 Bets on Agents
AI News

GPT-5.5: OpenAI's First Full Retrain Since GPT-4.5 Bets on Agents

OpenAI's GPT-5.5 is a fully retrained model, focusing on agentic computer use, not just benchmarks.

By Sarah Chen · 5 min · May 5, 2026

Mistral Medium 3.5: 128B Open-Weight Model That Opens PRs
AI News

Mistral Medium 3.5: 128B Open-Weight Model That Opens PRs

Mistral Medium 3.5 is a powerful 128B open-weight model capable of opening GitHub pull requests.

By Sarah Chen · 7 min · May 4, 2026

Microsoft Agent 365: $15-Per-Seat Control Plane for Your AI Agents
AI News

Microsoft Agent 365: $15-Per-Seat Control Plane for Your AI Agents

Microsoft Agent 365 offers a control plane to observe, govern, and secure all your AI agents.

By Sarah Chen · 6 min · May 2, 2026

DeepSeek V4 Pro: 1.6T Open-Weights Model Hits #2 on the Index
AI News

DeepSeek V4 Pro: 1.6T Open-Weights Model Hits #2 on the Index

DeepSeek V4 Pro is a top 1.6T open-weights model for agents, but has a high hallucination rate.

By Sarah Chen · 5 min · Apr 29, 2026

Coinbase's Fred and Balaji AI Agents Arrive in Slack
AI News

Coinbase's Fred and Balaji AI Agents Arrive in Slack

Coinbase launched AI agents modeled on Fred Ehrsam and Balaji Srinivasan in Slack and email.

By Sarah Chen · 5 min · Apr 21, 2026

OpenAI Agents SDK: Sandboxes Land for Long-Horizon Agents
AI News

OpenAI Agents SDK: Sandboxes Land for Long-Horizon Agents

OpenAI's Agents SDK now features sandboxes, built-in providers, and durable state for long-horizon agents.

By Sarah Chen · 5 min · Apr 20, 2026

Claude Opus 4.7: Anthropic's New Flagship Clears SWE-Bench Pro
AI News

Claude Opus 4.7: Anthropic's New Flagship Clears SWE-Bench Pro

Anthropic's Claude Opus 4.7 excels on SWE-bench Pro with enhanced vision and new features.

By Sarah Chen · 6 min · Apr 19, 2026

MAI-Transcribe-1: Microsoft's Whisper Killer Hits 3.8% WER at $0.36/Hour
AI News

MAI-Transcribe-1: Microsoft's Whisper Killer Hits 3.8% WER at $0.36/Hour

Microsoft's MAI-Transcribe-1 beats Whisper with 3.8% WER and lower costs, signaling independence from OpenAI.

By Sarah Chen · 6 min · Apr 17, 2026

Qwen 3.6 Plus: Alibaba's Free Preview Beats Claude Opus on Agent Tasks
AI News

Qwen 3.6 Plus: Alibaba's Free Preview Beats Claude Opus on Agent Tasks

Alibaba's Qwen 3.6 Plus Preview surpasses Claude Opus on agent tasks with impressive speed and context.

By Sarah Chen · 5 min · Apr 15, 2026

Figma for Agents: AI Now Designs Directly on Your Canvas
AI News

Figma for Agents: AI Now Designs Directly on Your Canvas

Figma now enables AI agents to design and modify directly on its canvas, leveraging your design system.

By Sarah Chen · 4 min · Apr 15, 2026

Atlassian Remix: AI Visuals and MCP Agents Come to Confluence
AI News

Atlassian Remix: AI Visuals and MCP Agents Come to Confluence

Atlassian Remix brings AI visuals and MCP agents to Confluence, transforming pages into dynamic content.

By Sarah Chen · 4 min · Apr 10, 2026

Meta Muse Spark: The First Model From Superintelligence Labs Is a Strategic Reset
AI News

Meta Muse Spark: The First Model From Superintelligence Labs Is a Strategic Reset

Meta Muse Spark, from Superintelligence Labs, marks a strategic AI reset with top benchmarks and medical reasoning.

By Sarah Chen · 5 min · Apr 9, 2026

Bluesky Attie: The AI Feed Builder That 125,000 Users Blocked on Sight
AI News

Bluesky Attie: The AI Feed Builder That 125,000 Users Blocked on Sight

Bluesky's Attie AI feed builder, powered by Claude, was blocked by 125,000 users quickly.

By Sarah Chen · 4 min · Apr 8, 2026

MindsDB Anton: The Open-Source BI Agent That Replaces Your Dashboard
Reviews

MindsDB Anton: The Open-Source BI Agent That Replaces Your Dashboard

MindsDB Anton is an open-source BI agent that replaces dashboards by answering questions in plain English.

By Sarah Chen · 4 min · Apr 8, 2026

Denovo Turns a Business Idea Into a Running Startup in 8 Minutes
AI News

Denovo Turns a Business Idea Into a Running Startup in 8 Minutes

Denovo's AI platform turns a business idea into a fully running startup in just eight minutes.

By Sarah Chen · 5 min · Apr 3, 2026

GLM-5V-Turbo: Z.ai's 744B Vision Model Turns Screenshots Into Code
AI News

GLM-5V-Turbo: Z.ai's 744B Vision Model Turns Screenshots Into Code

Z.ai's GLM-5V-Turbo vision model converts screenshots directly into executable code efficiently.

By Sarah Chen · 4 min · Apr 3, 2026

Tobira.ai: The AI Agent Network Where Bots Find You Business
AI News

Tobira.ai: The AI Agent Network Where Bots Find You Business

Tobira.ai is an AI agent network where bots find clients, partners, and investors for you.

By Sarah Chen · 5 min · Apr 2, 2026

Google Stitch 2.0: The Free AI Design Tool That Topped Product Hunt
AI News

Google Stitch 2.0: The Free AI Design Tool That Topped Product Hunt

Google Stitch 2.0, a free AI design tool, topped Product Hunt with new vibe design and voice canvas.

By Sarah Chen · 4 min · Apr 2, 2026

LillyPod: Eli Lilly's 9,000-Petaflop Supercomputer Bets Big on AI Drug Discovery
AI News

LillyPod: Eli Lilly's 9,000-Petaflop Supercomputer Bets Big on AI Drug Discovery

Eli Lilly's LillyPod, a 9,000-petaflop AI supercomputer, is making big bets on drug discovery.

By Sarah Chen · 4 min · Apr 1, 2026

NVIDIA Nemotron 3 Super: The Hybrid Architecture That Rewrites the Agent Playbook
AI News

NVIDIA Nemotron 3 Super: The Hybrid Architecture That Rewrites the Agent Playbook

NVIDIA's Nemotron 3 Super, a hybrid architecture, delivers 5x throughput and top agentic benchmarks.

By Sarah Chen · 4 min · Mar 31, 2026

Qwen 3.5 Small: Alibaba's 9B Model That Beats GPT-OSS-120B
Open Source

Qwen 3.5 Small: Alibaba's 9B Model That Beats GPT-OSS-120B

Alibaba's Qwen 3.5 Small, a 9B multimodal AI, surprisingly beats models 13x its size.

By Sarah Chen · 5 min · Mar 29, 2026

GPT-5.4: OpenAI's Five-Variant Strategy Reshapes the AI Market
AI News

GPT-5.4: OpenAI's Five-Variant Strategy Reshapes the AI Market

OpenAI's GPT-5.4, with five variants and expert-level computer use, is reshaping the AI market.

By Sarah Chen · 5 min · Mar 29, 2026