InfoQ AIGitHub
Android Bench 2 Adds Support for Long-Horizon Tasks, Agentic Evaluation, and Continuous Scoring

Google has released Android Bench 2.0, a major update to its benchmark framework for evaluating AI models and agents on Android development tasks.
Read at InfoQ AI ↗Related

GitHubPresentation: Building Reusable Evaluation Frameworks for Agentic AI Products InfoQ AI

GitHubPresentation: Multi-Agent Patterns from Spotify’s AI Powered Advertising Platform InfoQ AI
ValleyAI agent developer Manus raises $500M+ at reported $4B valuation SiliconANGLE
ValleyRein Security raises $25M to rein in insecure AI agents SiliconANGLE

ResearchEmotional optimization by newsroom AI needs behavioural evaluation Nature Machine Intelligence
