Videos

How to Monitor Infrastructure Health & Performance Using Randoli

September 25, 20256:49
How To with Randoli

A walkthrough of Randoli's Infrastructure Monitoring: real-time visibility into system health, host- and application-level performance tracking, and proactive issue detection, all from a single unified view.

Transcript

Hi, Kunal here, de engineer at Randoli. In this particular demo, I'm quickly going to walk you through our infrastructure monitoring capabilities designed to help DevOps and platform teams monitor system health, reduce performance degradations, and resolve issues faster across hybrid and multiloud environments all from a single view. Let's get started. What you see here is the infrastructure overview, which is a real-time consolidated view of your entire infrastructurees health across hybrid or multiloud environments. We begin by showing you a consolidated view of your host's availability across your entire infrastructure. Each tile here represents a host color-coded based on the availability. Green here means that the host is healthy and it's available. Red denotes that either it's unavailable or it's a failed host. Along with that we also surface the information on whether the host is on demand or is a spot instance which can help you to correlate any instability with that particular instance type. Just below this we have the issue trend by Kubernetes cluster and this can help you to quickly identify which particular cluster is showing signs of instability. This section uses severity based color coding which can help you to quickly identify which cluster may need further investigation. Clicking on a particular tile will reveal more cluster level insights. for example, resource usage, cluster level events, visibility into the network traffic across different services inside your cluster, workload related issues for that particular cluster and much more. We'll not go into the details of the cluster view. That's covered in a separate Kubernetes monitoring demo which you can surely check out. Now, monitoring resource health is critical to ensuring overall infrastructure stability. This particular section provides you at a glance view of the CPU, memory, and the disk usage of individual hosts. again using intuitive color coding based on that particular resource usage which can help you to quickly identify high pressure points or underutilized hosts. Now what makes this view more powerful is the ondemand telemetry snapshot that you can trigger for a particular host instead of continuously ingesting host level telemetry 24/7 which can be noisy and expensive at the same time. You can capture these detailed performance snapshots for a particular host within a dedicated time window. For example, let us say you wish to investigate a CPU spike that occurred yesterday between 2 and 3 p.m. You can go ahead and select that specific date and time and view the host telemetry during that dedicated time window. This will allow for a detailed timebound troubleshooting while also keeping your observability cost in check. Now, at the top of the telemetry snapshot, you'll find this metadata section which provides you some key system level metrics. for example, the host level details, system or OS level details such as the OS version, the architecture being used, and the system uptime as well. Now, depending on the type of host that you're viewing, this metadata section will adjust slightly. For example, in this particular case, the host is a Kubernetes node. So, you'll also find some cluster specific information. For example, the name of the cluster, the number of pods inside that cluster, the version of the cublet, and the condition for that particular cluster. You can even expand this to view more details for that particular cluster. Let's take a quick look at how this will differ if you are viewing the telemetry snapshot of a VM based host. As you can see, this is a telemetry snapshot of a virtual machine. And in the metadata section, you'll notice that only the host and the system or the OS level information is shown, which makes sense since this host is not a part of a Kubernetes cluster. Now, apart from the system level metadata, you can also view the resource usage of that particular host. for example, CPU, memory, and the disk usage. And here you have some quick visualizations of different but important metrics from a host perspective. For example, CPU and memory usage, the system load versus CPU cores, which you can filter out by the last 1 minute, last 5 minutes or the last 15 minutes just by clicking on these particular labels. As you can see, you can view the traffic related metrics such as network RX and TX, TCP connections, network RX and TX errors. Lastly, you can also view some disk related metrics such as disk read throughput, disk IOPS, disk read latency as well. So overall, this gives you a detailed in-depth view of how a particular host is performing during a specific time window or while you're investigating a particular issue. Now, in addition to viewing the CPU and the memory usage of that particular host, you have the ability to view the resource usage over time and its commitment. This is especially helpful when a host is overcommitted. Meaning the total workloads which are scheduled exceeds its actual capacity. That's a red flag for potential throttling or failures which you can quickly identify from this particular view. As an engineer, this can help you to make informed scheduling decisions. Should we rebalance the workloads across host? Should we scale out the capacity or should we clean out some unused services? These commitment metrics can directly impact such operational decisions. So at the end of the day we can aim to maintain the performance and reduce and prevent saturation at the same time. All right. Now let's talk about proactive detection. What you see here are timeline based visualizations of the triggered monitors from the host level as well as the application level. Now with these monitor timelines in the infrastructure overview, you have the ability to click on any particular data point and view all the triggered monitors for that specific time window. These alerts can be filtered based on the severity and comes with all the detailed contextual information that you may need in order to debug that issue further. For instance, the name of the cluster if that particular host is a Kubernetes node, the severity of that particular alert and a short description which tells you why this particular alert was triggered. In addition to this, you also have the ability to quickly navigate to that particular host's telemetry snapshot where you can quickly analyze how that particular host was performing during that specific time window when the alert was triggered. This kind of correlation can help you to enhance your overall debugging process so you can resolve your issues much faster. And the same goes for the application level monitors as well as you can see right here. So overall this enables proactive detection across infrastructure as well as the application level helping your team to detect and correlate issues much faster and taking targeted actions before these escalate and affect your users. So that's how Randoli helps you to monitor your entire infrastructure which could be spanned across a multi cloud or a hybrid environment all from a single view. It gives you a centralized cost-effective way to track availability, detect issues proactively, and troubleshoot based on actual performance patterns. If you'd like to explore this further or see how it may fit your environment, feel free to check out the documentation or reach out to us. We would love to hear about your use case. Thank you for watching.

See how Randoli applies this in practice.