/

Apple Machine Learning Research

9 stories

Apple Machine Learning Research

Apple Machine Learning ResearchResearchREFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs Most current vision-language-action (VLA) models—such as OpenVLA, π0, RT-2, and RDT-1B—are “monolithic.” This means they generate raw motor commands or very short sequences of actions, without organizing behaviors into r

September 2
Apple Machine Learning Research

Apple Machine Learning ResearchResearchLLMs Are Not (Consistently) Bayesian: Quantifying Internal (In)consistencies of LLMs’ Probabilistic Beliefs Modern AI systems are being deployed in complex domains such as medicine, science, and law, where there is often not a single correct answer given the observed evidence. Such systems must be able to represent and update

August 28
Apple Machine Learning Research

Apple Machine Learning ResearchResearchAgent Seer: Synthesizing Scenarios from Specification Understanding Evaluating AI agents that use external tools requires realistic test scenarios that capture how practitioners compose tools and iterate across conversation turns. Constructing such scenarios by hand demands deep domain e

August 28
Apple Machine Learning Research

Apple Machine Learning ResearchResearchFrom Preferences to Principles: Rubric-Based Alignment for Grounded Knowledge Answers Designing effective reward signals for open-domain question answering is challenging because high-quality responses must simultaneously satisfy multiple aspects of answer quality that are difficult to capture with a holi

August 27
Apple Machine Learning Research

Apple Machine Learning ResearchResearchPROOF-Gen: From Optimized Data to Better Distillation Supervised fine-tuning on teacher-generated trajectories is the standard first stage for distilling tool-calling capabilities into deployable models. Post-training pipelines that drive shipped tool-calling agents re-run

August 26
Apple Machine Learning Research

Apple Machine Learning ResearchResearchLuce: Relightable Gaussians for 3D Asset Generation High-fidelity image-to-3D generation requires a 3D representation that captures both geometry and appearance. To support relighting and integration into standard rendering pipelines, the representation should include phy

August 26
Apple Machine Learning Research

Apple Machine Learning ResearchResearchSTARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation Unified multimodal models that understand, reason over, and generate interleaved text–image sequences remain structurally fragmented: existing approaches either sacrifice visual fidelity through discrete tokenization, im

August 25
Apple Machine Learning Research

Apple Machine Learning ResearchResearchBeyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning Multimodal large language models increasingly use visual chain-of-thought (Visual CoT) to reason about spatial, temporal, and embodied environments. By generating intermediate reasoning images, Visual CoT provides an int

August 24
Apple Machine Learning Research

Apple Machine Learning ResearchResearchMultilingual Knowledge Transfer under Data Constraints via Lexical Interventions Cross-lingual knowledge transfer is critical for building high-performing multilingual language models for languages with insufficient training data. When target language data is scarce, the knowledge required for many d

August 20