Gremlin’s cover photo
Gremlin

Gremlin

Software Development

San Jose, California 12,365 followers

The Reliability Management Platform for high-velocity engineering teams

About us

Gremlin’s Reliability Management Platform enables high-velocity engineering teams to standardize and automate reliability across their organizations without slowing down software delivery. Gremlin's Reliability Score sets the standard for reliability so there's no guesswork, and an automated suite of Reliability Management tools makes it easy to integrate reliability throughout the software lifecycle so there's no slowdown.

Website
http://www.gremlin.com
Industry
Software Development
Company size
51-200 employees
Headquarters
San Jose, California
Type
Privately Held
Founded
2016
Specialties
Distributed Systems, Resilience, Failures as a Service, DevOps, and Chaos Engineering

Locations

Employees at Gremlin

Updates

  • Reliability only works if the team actually uses the tooling. It sounds obvious, but it's one of the most common ways reliability initiatives stall. Engineering teams have enough on their plates. If deployment is manual, if onboarding is painful, if tests have to be configured from scratch every time- people stop using it. Look for automated deployment, centralized admin controls, and automatic service catalog creation. The faster teams can get up and running, the faster your reliability practice grows. Find the latest recommendations in our 2026 Chaos Engineering buyer’s guide. ⬇️ https://hubs.la/Q04drxwM0

  • AI-generated code has 1.7x more issues per PR than human-written code... and teams are shipping 10x the PRs. Reliability guardrails don't just catch that at the gate. Fed back as context, resilience test results help AI coding agents produce better code over time- the same way engineers get better at avoiding failure once they've seen it happen. They also cut resolution time for AI SREs. If 6 of 8 microservices already passed resource scaling tests, an outage investigation can skip straight to the 2 that didn't. The trick isn't choosing between speed and reliability.Iit's automation that gives you both. Read more at the link in the comments.

    • No alternative text description for this image
  • Gremlin reposted this

    I had the opportunity to sit down with Dana Kohut at The Prime View recently to discuss Reliability in the AI ERA. We had a great conversation! A few things we discussed: AI is writing more code than ever. Is it also writing better code? What do “AI guardrails” actually mean in practice, not as a buzzword, but as something you build? Customers used to be nervous about letting outside systems touch production. Has that changed with AI? With AWS, Azure, and Google all now offering some version of chaos and reliability testing, why does a dedicated guardrail layer like Gremlin still matter? Does this mean SREs eventually become unnecessary – fully self-healing systems, no humans required? Take a moment to read it, or give it a listen! Let me know what you think I got wrong!

  • Gremlin reposted this

    People think 4-nines of uptime are out of their grasp, but it is very achievable by doing the basics. Most companies operate between two and three nines of availability. That’s 8 hours to 3 and a half days of downtime a year. Many organizations don’t want to talk about their outages because it makes them look bad. But we’ve all had them, and the way to improving is to start taking action today. It’s easier than you think. You can achieve four 9s just by doing the basics. Testing what happens if you lose a host or an availability zone. Knowing each of your dependencies (internal and external) and testing what happens if they stop responding. These are the core tests behind any reliability testing effort. And you’d be surprised by how much of a difference they can make.

  • Reliability guardrails help make sure that your applications stay reliable without slowing down. But that’s only the beginning. By themselves, guardrails act as a gate to ensure resilience mechanisms hold under rapid changes. When set up to create a feedback loop, they can also help AI agents produce higher-quality code and react faster when an outage does occur. Keep reading at the link in the comments to learn more...

    • No alternative text description for this image
  • Engineering teams are under pressure to move faster. At the same time, expectations for uptime and resilience keep increasing. That creates a problem: Reliability testing often becomes too slow, too manual, or too difficult to implement consistently. Failure Flags by proxy reduces that friction: ➡️ No code changes ➡️ Lightweight deployment ➡️ Built-in health checks ➡️ Application-level failure injection So teams can spend less time coordinating tests and more time improving reliability. Read more at https://hubs.la/Q04k6GQ40 

  • Excited to see Gremlin included in SD Times Top 100 list 🎉 "...the companies in this year’s Continuous Quality & Validation category are largely defined by how they’re using AI and automation to close that widening gap rather than simply asking teams to test faster with the same manual effort." Music to our ears! Testing has to evolve as the AI landscape does- and our goal is to make this even faster and easier for Gremlin customers. https://lnkd.in/etF_5emA

  • Chaos Engineering is supposed to find failures before they find you. 👀 But picking the wrong tool, or one that isn’t prepared to deal with the influx of risk in an increasingly AI-driven world, is a recipe for disaster. The 2026 Chaos Engineering Buyer's Guide breaks down exactly what to look for- link in the comments.

  • Traditional observability is reactive by design. It shows you what already failed. Failure Flags helps teams move toward predictive reliability: 🚀 Validate failure modes proactively 🚀 Measure resilience before incidents occur 🚀 Surface hidden risks traditional monitoring misses And now, with Failure Flags by proxy, teams can start testing serverless applications without modifying application code. Get started at https://hubs.la/Q04k726N0

  • 🗓️Tomorrow: More Resilience, Less Overhead: How to Modernize Disaster Recovery Testing We’re leading a live session covering: ➡️ Where current disaster recovery verifications fall short and leave blind spots ➡️ How to break down the most common catastrophic failures into systematic individual tests ➡️How to accurately simulate disaster scenarios organization-wide with a fraction of the lift Last chance to register: https://hubs.la/Q04jKty_0

Similar pages

Browse jobs