The New Stack
24 stories

The New StackGitHubClaude Fable 5.1 vs. Fable 5: On real work, I couldn’t tell them apart. Anthropic launched Claude Fable 5.1 this week, calling it “our most advanced model for coding and knowledge work.” There was

The New StackGitHubBuilding trust in agentic RAG starts with evidence Basic retrieval-augmented generation (RAG) follows a straightforward pattern. A user asks a question, the system finds relevant content in a

The New StackGitHub“Sorry for the messy rollout”: OpenAI launches GPT-6 Astra to most paying users a day after its unveiling Update: As of 6:30 p.m. Eastern on Friday, September 4, GPT-6 Astra was available on ChatGPT for all paying users

The New StackGitHubMicrosoft built a prompt injection detector. Then it caught a phishing campaign instead. Microsoft flagged a phishing campaign last week that exploits a gap in how machines read text. Attackers are slipping invisible

The New StackGitHubOpenAI will sell you Astra, but not the system that scored 98.6% on ARC-AGI-3 Investor Matt Turck, whose fantastic podcast has hosted the people who built ARC-AGI, summed up Astra’s blockbuster benchmarks with three

The New StackGitHubAI agent evaluations are part of the product A team builds an agent, gives it a few representative questions in a test chat, and watches it produce useful

The New StackGitHub“1% of my engineers are responsible for 40% of token spend”: Why Coder and SpaceXAI want to give developers nice things Coder announced its Coder Agent Relay service this week, with SpaceXAI as its launch partner. The service lets software engineering

The New StackGitHubOpenAI spends $1 billion to expand Daybreak to defend power, water, and banking In a livestreamed keynote on Thursday, OpenAI president Greg Brockman announced Daybreak for Frontline Defenders, a new global initiative to

The New StackGitHubGPT-6 Astra’s score of 98.6% looked like AGI. Then researchers read the fine print. There was no mistaking the divide in March with the release of ARC-AGI-3. While frontier AI models could do little

The New StackGitHubThe systems guide to production token optimization When enterprise AI applications scale, they inevitably hit a wall. For many engineering teams, this wall is initially diagnosed as

The New StackGitHubCut GPU inference cold start from 8 minutes to less than a minute We instrumented the full path from pod creation to first inference response on a GPU node running a 70B-class model.

The New StackGitHubOpenAI launches GPT-6 Astra and says welcome to the “AGI era” OpenAI on Thursday launched GPT-6 Astra, its newest flagship model. The company describes it as “the world’s most intelligent and

The New StackGitHub“Hugging Face will remain an open platform”: Nvidia strikes $12.9B deal for the ‘GitHub of AI’ Nvidia has confirmed that it’s agreed to acquire Hugging Face in a mammoth $12.9 billion deal that will bring one

The New StackGitHubAI Agents built a 3D city for $33 in two hours —and exposed a major flaw PhiloLabs wanted to see how far a group of AI coding agents could get building something where working code wasn’t

The New StackGitHubNvidia PAIR lets you put your idle Macs and PCs to work for AI agents Nvidia’s bet on open models and local AI has been taking shape for a while now. Its acquisition of Hugging

The New StackGitHubWant to scale AI agents without breaking anything? Retrieval engineering is the answer. AI agents are multiplying as corporations adopt the technology in record numbers. Smarter underlying models, better tool use, and improved

The New StackGitHub“Google was ahead only a few hours”: Muse Spark 1.3 edges out Gemini as Meta claims its biggest coding leap yet It has been a big week for Meta in the AI sphere, formally launching its Muse Code coding agent out

The New StackGitHubMultiverse says its 438B model is fast enough for AI agents. The benchmarks tell a more complicated story. A 438-billion-parameter reasoning model isn’t an obvious choice when speed is a priority. Multiverse Computing is betting that compression can

The New StackGitHubYour next OpenAI API timeout might not be a timeout at all OpenAI said Tuesday that its upcoming Astra model is the company’s first to reach the Critical cybersecurity threshold in its

The New StackGitHubAnthropic’s Claude failures have made agent observability a security priority Anthropic aimed to steer its ship into safer, more carefully charted waters this week. The company announced it was improving

The New StackGitHubGoogle ships its third Gemini Flash model in six weeks Google hasn’t released a Gemini Pro model in a while, but on Wednesday, the company launched yet another set of

The New StackGitHubVercel built a feedback loop that treats agent instructions like software Vercel ran more than 200 agent runs to build design.md, a new public prompt file designed to help agents create

The New StackGitHubYour organization prioritized AI adoption, but you actually need AI fluency. Thanks to increasingly capable models, some parts of your business are getting faster, more capable, and more productive every month.

The New StackGitHubClaude Fable 5.1 watermark: It has a blind spot developers can’t ignore Anthropic launched Claude Fable 5.1 on Tuesday with a statistical signature embedded in its generated text, but developers shouldn’t expect
Nothing matches this filter yet.