/

研究

126件

Scaling Agentic RL: High-Throughput Agentic Training with Tunix

Google Developers BlogLabsScaling Agentic RL: High-Throughput Agentic Training with Tunix Tunix is Google’s new JAX-native post-training library designed to eliminate TPU idling bottlenecks when training multi-turn, tool-using LLM reasoning agents. It maximizes hardware throughput by combining highly concurre

9月6日
How to use Google microbenchmarks for evaluating TPU performance

Google Developers BlogLabsHow to use Google microbenchmarks for evaluating TPU performance Google's open-source TPU microbenchmark suite provides developers with granular performance metrics across Network, Compute, HBM, Host Transfer, and Attention components to validate real-world hardware capabilities. By l

9月6日
Agent and Model Evaluations in Gemini Enterprise Agent Platform are now GA

Google Developers BlogLabsAgent and Model Evaluations in Gemini Enterprise Agent Platform are now GA Agent Platform's evaluation service is now generally available, providing developers with a unified engine to measure agent quality consistently across local development experiments and live production traffic. You can e

9月6日
Scaling real-time AI agents with session-aware load balancing

Google Developers BlogLabsScaling real-time AI agents with session-aware load balancing Real-time AI agents break traditional request-response load balancing paradigms because they rely on long-lived, stateful bidirectional streams that obscure true server capacity. To solve this, developers must implement

9月6日
Scaling AI Agent Infrastructure with the MCP Stateless updates

Google Developers BlogLabsScaling AI Agent Infrastructure with the MCP Stateless updates The 2026-07-28 Model Context Protocol (MCP) specification replaces legacy stateful constraints with a fully stateless core, enabling cloud-native horizontal scaling, serverless deployments, and standard round-robin load

9月6日
How to Evaluate Live & Voice Agents in ADK

Google Developers BlogLabsHow to Evaluate Live & Voice Agents in ADK Moving live voice agents from demo to production requires rigorous, automated testing to handle the unpredictability of real multi-turn conversations. ADK now provides native live evaluation, allowing developers to test

9月6日
Zenn (AI)

Zenn (AI)JapanAIエージェントの評価を実装する:本番ドリフトを検出するEvals設計 ※本稿は、当方ブログの紹介文です。全文はこちらへどうぞ。 https://shinichi.noguchi.jp.net/blog/2026-09-05-agent-evals-in-production.html AIエージェントの評価を、単一の正解率ではなく、実行トレース・生成品質・本番挙動の3層へ分解します。評価対象を分解しないまま「回答が悪くなった」とだけ議論すると、プロンプト、ツール、モデル、入力分布のどこに回帰があるのか切り

9月5日
Beyond Zero: Google Publishes Successor to BeyondCorp

InfoQ AIGitHubBeyond Zero: Google Publishes Successor to BeyondCorp In a recent research paper, Google introduced Beyond Zero, a “security model for the AI era” that extends Zero Trust to autonomous AI agents. The new approach moves access decisions from the application level to individu

9月5日
The Pelican comparison grid for Astra is pretty interesting

Simon WillisonVoicesThe Pelican comparison grid for Astra is pretty interesting I got access to GPT-6 Astra this afternoon, so naturally I used it to generate SVGs of pelicans riding bicycles - at low, medium, high, xhigh and max reasoning levels (Astra doesn't support reasoning=none). Then I render

9月4日
Frontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson

NVIDIA Technical BlogSiliconFrontier Reasoning Reaches the Edge: How to Deploy and Optimize Models on NVIDIA Jetson Running reasoning and agentic AI at the edge has been harder than it needs to be. Until recently, models capable of multi-step reasoning were too large to run...

9月4日
Project HydraFusion: Frontier quality via multi-model orchestration

GitHub BlogGitHubProject HydraFusion: Frontier quality via multi-model orchestration In controlled offline evaluations, HydraFusion’s selective coding workflows matched or exceeded the evaluated Opus 5 baseline while reducing estimated workflow cost. Now available as a research preview in GitHub Copilot.

9月4日
AI agent evaluations are part of the product

The New StackGitHubAI agent evaluations are part of the product A team builds an agent, gives it a few representative questions in a test chat, and watches it produce useful

9月4日
Mini book: Next-Gen Architecture Playbook: Insights and Patterns for the AI Era

InfoQ AIGitHubMini book: Next-Gen Architecture Playbook: Insights and Patterns for the AI Era This eMag examines how architects can lead with clarity in a rapidly evolving engineering world, distilling industry insights into field-tested practices for teams. Together, these stories reveal a core theme: the techno

9月4日
Stratechery

StratecheryVoicesAn Interview with OpenAI President Greg Brockman About Astra and Alignment An interview with OpenAI President and Co-Founder Greg Brockman about the history of OpenAI, Astra and alignment, and the weight of building the future.

9月4日
Want to scale AI agents without breaking anything? Retrieval engineering is the answer.

The New StackGitHubWant to scale AI agents without breaking anything? Retrieval engineering is the answer. AI agents are multiplying as corporations adopt the technology in record numbers. Smarter underlying models, better tool use, and improved

9月3日
Scaling agentic AI pilots across the enterprise

MIT Technology ReviewValleyScaling agentic AI pilots across the enterprise As agentic AI moves from experimentation toward enterprise deployment, the challenge is figuring out how agents can work together, connect to the systems and data they need, and operate safely across the workflows that r

9月3日
Practical AI

Practical AIVoicesLess about Models; More about Architecture As AI moves from experimentation to enterprise deployment, are organizations thinking too much about models and not enough about architecture? In this episode, Daniel and Chris talk with Chetan Gupta, Chief AI Officer at

9月3日
Software Engineering Daily

Software Engineering DailyVoicesMoving Beyond RAG with Precomputed Context Retrieval has become one of the central problems in building useful AI systems. The standard approach to grounding a model in one’s own data has been retrieval augmented generation, or RAG, where an agent searches a vect

9月3日
Claude Fable AI Is Much Stranger Than The Headlines Suggest

Two Minute PapersResearchClaude Fable AI Is Much Stranger Than The Headlines Suggest ❤️ Check out Lambda here and sign up for their GPU Cloud: https://lambda.ai/papers 📝 The Claude Fable 5.1 paper is available here: https://www.anthropic.com/claude-fable-and-mythos-5-1 https://www-cdn.anthropic.com/0339e

9月3日
Multiverse says its 438B model is fast enough for AI agents. The benchmarks tell a more complicated story.

The New StackGitHubMultiverse says its 438B model is fast enough for AI agents. The benchmarks tell a more complicated story. A 438-billion-parameter reasoning model isn’t an obvious choice when speed is a priority. Multiverse Computing is betting that compression can

9月2日
Modernizing and scaling support operations with generative AI on AWS

AWS Machine LearningLabsModernizing and scaling support operations with generative AI on AWS Learn how to build a generative AI-based support operations platform on AWS that converts training videos into structured SOPs, applies Retrieval-Augmented Generation to guide ticket resolution, and uses machine learning

9月2日
From code to diagrams: Agentic architecture documentation with Amazon Bedrock AgentCore

AWS Machine LearningLabsFrom code to diagrams: Agentic architecture documentation with Amazon Bedrock AgentCore Learn how a global interdealer broker built an automated architecture documentation pipeline on Amazon Bedrock AgentCore that analyzes .NET code bases, generates architecture diagrams, and maintains searchable documentat

9月2日
Hacker News Show

Hacker News ShowGitHubShow HN: FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x

9月2日
Presentation: Beyond Prompting: Context Engineering for Production-Grade AI

InfoQ AIGitHubPresentation: Beyond Prompting: Context Engineering for Production-Grade AI Ricardo Ferreira discusses moving beyond simple prompt engineering to build production-grade AI applications. He shares practical architectural strategies for integrating long-term and short-term memory using Redis, mana

9月2日