RED Metrics
Randoli's CTO explains RED metrics, rate, errors, and duration, and why they surface problems that CPU and memory monitoring alone can miss.
Transcript
Hello, Rajat here. Today I want to talk to you about red metrics and why it is important to collect and monitor red metrics. Red metric stands for rate, errors and duration. Rate is how often your requests happen, also known as throughput. Errors, how often they fail. and duration how long they take also known as latency. Why is it important to measure uh or collect and uh you know monitor red metrics? Um CPU and memory most people are monitoring CPU and memory. Uh they are a great way to understand the health of your application but they not they may not necessarily u help you understand what your users are experiencing. For example, just because your CPU and memory um is not uh trending higher doesn't necessarily mean that your users um um are happy because they may be experienc a high number of errors. The requests may be taking longer to process and they don't necessarily um correlate to high CPU usage or high memory usage. So if you're just measuring CPU and memory, you may not uh get to know what your users are experiencing. And on the flip side, just because your CPU and memory is uh trending higher, it may not necessarily have an impact on your customers experience either. So how can it help you to monitor user experience? Let's take an order service as an example. And you have this goal where you want to ensure that you keep the errors uh to about 5% or less because there's always going to be some sort of error. So you have this objective of making sure that if it ex you know you want to keep it under 5% and if it exceeds 5% then I want to be notified because I want to take some action to investigate what's going on to rectify the situation before it impacts my customers experience. Um similarly you could say I want to ensure that my customers can place an order in 50 milliseconds or less. Anything more is going to impact uh my customers experience and they may not um use my service in the future. So you could have a uh alerting condition based off um the the duration metric and say my P99 which means 99% of my uh request will be served um you know less than 50 milliseconds. So they have you know one person could exceed your P99 uh but that's okay. But I want 99% of my request to not exceed 50 milliseconds and I want an alert so I can uh be notified if you know my system is not doing that so I can and investigate and figure out what's what's going on. Similarly, uh this can extend to um you know database queries, uh looking up something from a radius cache, um executing a query, uh by measuring individual actions that can actually indicate trouble. So for example, let's say you have a couple of database queries. We all have queries uh that are that could be troublemakers, right? So you want to keep an eye on them. Let's say you want a database query to ex uh you know execute within um you know 10 milliseconds. If it exceeds that you know it's trending towards trouble right or if your database query is uh throwing a lot of errors that could indicate um some other issue. So you want to be on top of these issues and the only way you can do that is if you have red metrics rate errors and duration uh for these individual uh operations and and and then you can also take those operations and aggregate uh at the service level and if you have the red metrics you could do that and then you could also take multiple services that are part of a user journey and aggregate them um to that next level and say hey I want my entire order flow to have less than 5% errors or I want my entire order flow to execute within uh you know like 50 milliseconds or 100 milliseconds uh to help make sure that I am serving my customers the way I should be. So red metrics is a great way to measure custom experience. It's also the building blocks for you to build a more comprehensive SLO monitoring feature. Now that we talked about the importance of SLOs's, let's look at um how we can u start collecting red metrics. Open telemetry great place to start. So open telemetry traces um has something called span metrics which is actually computed based off the uh the spans in your uh open traces. It measures the duration, errors and requests from your spans um and creates metrics out of it uh which you can uh store it in Prometheus or uh send it to your observability vendor and then you can use those to build uh all kinds of monitoring analysts that I talked about. So when you are using open telemetry tracers, make sure you uh enable uh span metrics in your open telemetry pipeline. Once you have that, you could start to do uh very granular level monitoring. you can roll up the metrics to do um SLOs's at the service level and then keep rolling up further uh to monitor SLOs's at the user journey level and you know and as a whole uh by aggregating those metrics. Hope you find this um video useful and if you have any questions or comments please feel free to uh you know let us know and uh we hope to see you uh in a further video and I'll explain um how to configure open telemetry pipeline how to enable um span metrics etc. See you. Thank you.