Modal
25 stories

ModalToolsHitting a billion tokens per minute on one GPU by combining a query planner and an inference engine Maximizing perf on AI-SQL queries with the KV-optimal left-deep join

ModalToolsHow to serve trillions of tokens for trillion-parameter coding agents Learn how we optimized performance and efficiency serving the workload that is changing software engineering forever — and you can too.

ModalToolsProduct updates: Sandbox Sidecars, new models, a refreshed dashboard, and more Recent product updates from Modal and news from around the community.

ModalToolsModal is expanding in Europe with our new London office Modal is expanding, and hiring on all fronts across Europe.

ModalToolsHow Botika runs full-stack generative AI on Modal Botika is building the next generation of tooling for agentic e-commerce brands, backed by a fleet of custom models trained and served on Modal.

ModalToolsQwen3.8-2.4T-A95B now available on Modal Qwen3.8-2.4T-A95B by Alibaba, with a 1M token context window, is now available via Modal Auto Endpoints.

ModalToolsBringing serverless functions closer to the speed of wire Modal’s Function Call data path is now >50ms faster. Our new routing layer is geographically distributed, so you can further reduce your network overhead.

ModalToolsA note on the Hugging Face agent incident Hugging Face published a technical timeline of a recent agent intrusion. Modal's platform and isolation were not compromised in this incident.

ModalToolsKimi K3 by Moonshot now available on Modal Kimi K3, a 2.8 trillion parameter multimodal model by Moonshot, along with a custom-trained DFlash speculator, is now available on Modal.

ModalToolsDevin Outposts on Modal Devin, built by Cognition, is an AI software engineer: it plans, writes, tests, and ships code semi-autonomously. With Outposts, Devin can now run its work in Modal sandboxes.

ModalToolsScaling to 1 million concurrent sandboxes in seconds How (and why) we built a scheduling system that can scale to 1 million concurrent sandboxes (per workspace) in seconds.

ModalToolsInkling by Thinking Machines now available on Modal Inkling, a general-purpose multimodal model by Thinking Machines, along with a custom trained DFlash speculator, is now available on Modal.

ModalToolsModal is a computer What is Modal? A machine that runs programs of arithmetic & logic operations on information.

ModalToolsHow to price serverless GPUs To compare rates for serverless and reserved GPUs, look at your application's peak-to-average ratio.

ModalToolsMulti-token Residual Prediction One tiny module, two ways to win on Diffusion LMs

ModalToolsAnthropic integration with Modal brings scalable compute to Claude Science Announcing our integration with Claude Science, bringing Modal's elastic compute to researchers when they need it.

ModalToolsModal's serverless Servers A deep dive inside our new ultra-low-latency primitive.

ModalToolsAchieve state-of-the-art inference latencies with speculative decoding How Modal and Decagon worked together to cut inference latency - and you can too.

ModalToolsIntroducing Modal Auto Endpoints: Optimized inference you actually own LLM inference at SotA speeds and Modal quality, now available to everyone.

ModalToolsUnpacking sandbox startup latency: why started ≠ ready Building performant sandbox systems goes way beyond the initial container boot. Here, we unpack what that means, and discuss some tools to help you manage the entire lifecycle.

ModalToolsSpeculation Is All You Need Why we're all-in on speculative decoding.

ModalToolsProduct updates: VM Sandboxes, Lower latency routing, RBAC, and more Recent product updates and news from around the community.

ModalToolsMaking FlashAttention-4 faster for inference What part of "dtype = 'fp8', num_splits = 0, pack_gqa = True, q_stage = 1, page_size = 1" do you not understand?

ModalToolsReinforcement learning is an infrastructure problem What we've seen helping teams run Reinforcement Learning at scale on Modal. Plus an open-source library to skip the scaffolding.
Nothing matches this filter yet.