Inkling by Thinking Machines Lab is now available on Modal, backed by a custom DFlash speculator for 67% higher throughput and interactivity. Try it out on Modal today. Read more on our blog: https://lnkd.in/gemkn4nT
Modal
Software Development
New York City, New York 27,342 followers
AI needs a new infrastructure layer. We're building it.
About us
Customers rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale. Every era of computing came with new workloads that previous infrastructure couldn't serve: mainframes, databases, the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice. The window to build is open right now.
- Website
-
https://modal.com
External link for Modal
- Industry
- Software Development
- Company size
- 51-200 employees
- Headquarters
- New York City, New York
- Type
- Privately Held
- Specialties
- Serverless GPUs, LLM Inference, LLM Fine-Tuning, Generative Model Inference, Generative Model Training, Computational Biology, Audio Generation, Image Generation, Video Generation, Web Scraping, Batch Jobs, Batch Embeddings, Scaling Out, AI Agents, Reinforcement Learning, Sandboxes, and Background Agents
Products
Modal
Platform as a Service (PaaS) Software
Modal is a serverless compute platform that makes it easy for developers to run compute-intensive workloads like ML inference, fine-tuning, and batch jobs. Our proprietary Rust-based container stack is best-in-class, allowing you to run any function in the cloud in less than a second, even on the most in-demand GPU types. We autoscale to thousands of GPUs or CPUs for your functions based on request volume so you can always meet customer demand while never paying for idle resources. Modal's Python SDK allows you to define custom images and hardware requirements in code. No more spending time on config files or cloud consoles. Let your team ship innovative AI products—we'll handle the compute.
Locations
-
Primary
Get directions
New York City, New York 10038, US
-
Get directions
Stockholm , SE
-
Get directions
San Francisco, California 94103, US
Employees at Modal
Updates
-
Modal reposted this
Bonjour, RAISE Summit 2026! 🇫🇷 We're in Paris July 7-9. Come find the Modal team at Booth 32B. Don't miss our CEO Erik Bernhardsson's fireside chat on Thursday at 9:40 CET (July 9), where he'll share his take on the future of cloud computing and what it takes to scale engineering culture at Modal. To close out the conference, pop by our champagne, caviar and nuggets gathering with Poolside, Gladia, turbopuffer and Vercel. See you soon!
-
-
Bonjour, RAISE Summit 2026! 🇫🇷 We're in Paris July 7-9. Come find the Modal team at Booth 32B. Don't miss our CEO Erik Bernhardsson's fireside chat on Thursday at 9:40 CET (July 9), where he'll share his take on the future of cloud computing and what it takes to scale engineering culture at Modal. To close out the conference, pop by our champagne, caviar and nuggets gathering with Poolside, Gladia, turbopuffer and Vercel. See you soon!
-
-
Modal is co-hosting Champagne, Caviar, AI & Nuggs at RAISE Summit with our friends from Poolside, Gladia, turbopuffer and Vercel. Come join us after day two of the conference and put your taste to the test on the luxurious, delicate, beloved chicken nuggets. July 9. Come say hi 👋
-
-
The most demanding problems in life sciences need more than a capable model, they need infrastructure that scales. Today we're announcing our integration with Claude Science, bringing Modal's elastic compute to researchers when they need it. We're committing up to $100K in compute to support academic life sciences research. Apply by July 15. Read more: https://lnkd.in/gf3X5T26
-
Modal reposted this
We OCR'd 100,000 pages with open-source vision models on Modal in under an hour for about $225. The same job with GPT-5.5 would have cost close to $6,000. The lower-cost proprietary models were cheaper, but they hallucinated enough that it just wasn’t a worthwhile comparison. We wanted to know what it would take to OCR a large document corpus using open-source vision models instead of paying for proprietary APIs. What surprised us the most was that the cheapest GPU per second wasn't the cheapest GPU for the job, in the end. L4s and H100s landed at roughly the same cost per page in some runs, but L4s took 5-6x longer. Timeline should be considered part of the cost. We’ve got a full write-up with methodology, model comparisons, and a side-by-side viewer to judge output quality yourself > https://lnkd.in/gv7VUyT8
-
-
Low-latency inference demands a new serving primitive: Servers. Modal Servers are designed for applications where every millisecond counts, like LLM inference for interactive agents. Servers give you a regionalized, autoscaling pool of HTTP server replicas behind Modal’s routing layer with the deployment ergonomics, fast feedback loops, and autoscaling you know and love. We get into the specifics of how we built this in our new post: https://lnkd.in/gmDFzzGS
-
Modal Auto Endpoints provide state-of-the-art inference performance out of the box. This is because each Endpoint is backed by a low latency inference playbook developed in concert with leading AI companies like Decagon, delivering responses 60ms faster than the best proprietary providers. Learn how: https://lnkd.in/g4Yfdcat
-