How logs can help to improve security and observability
How logs can help to improve security and observability
This webinar recaps how Chronosphere’s Telemetry Pipeline streamlines your log data management and optimizes your data footprint.
Sophie Kohler | Social Media and Content Manager | Chronosphere
Sophie Kohler is a Content Writer at Chronosphere where she writes blogs as well as creates videos and other educational content for a business-to-business audience. In her free time, you can find her at hot yoga, working on creative writing, or playing a game of pool.
On: Jan 24, 2025
21 MINS READ
Log data is skyrocketing
Log data has skyrocketed by 250% in just the past year alone. With sources multiplying, formats diversifying, and the number of destinations increasing, managing this data is becoming more costly and complex for teams.
In a recent webinar, Chronosphere Director of Product & Solutions Marketing, Riley Peronto, and Chronosphere Field Architect Anurag Gupta dive into how Chronosphere’s Telemetry Pipeline solution can streamline your log data management, optimize your data footprint, and alleviate the burden of telemetry collection and routing.
If you don’t have time to listen to the full chat, catch a transcript below.
Observability and security teams: What we’re seeing today
Riley Peronto: I’ll start off by talking a little bit about the Chronosphere Telemetry Pipeline and the challenges that we’re solving as an org. Then I’ll hand it off to Anurag to talk about why you might want pre-processing of your log data within your organization. Then, we’ll shift over to the actual demo. We’ve got use cases that are both basic and advanced, as well as some strategies to help you reduce log data inflate. With that, let’s talk a little bit about the challenges we solve as an organization.
When we’re talking to security and observability teams, we see two things across the board. First, I see that infrastructure is increasing in complexity. So, what do I mean by that? Teams are sending data to more backends than they ever have before. That might be a log management or an observability platform. It could be a SIEM platform.
You could even be sending data to a data warehouse or a data lake or something of that nature. And these tools, while there’s many of them, they’re adding a ton of value to security and observability practices. And they’re helping teams really get more value out of their logs. However, as you add more backends and parallel, you also need to add more ways to collect that log data and ship it downstream.
Management burden and battling complexity
This is where the complexity can start to creep in. What we see is that teams are adding more agents upstream than they’ve had before. And this makes it harder to enforce changes upstream. It makes it harder to configure everything in a consistent manner. It can make it hard to make sure that you’re getting all the logs you need without any blind spots.
Overall, this creates a big management burden that the engineering team or whoever is responsible for your logging infrastructure has to keep up with, and it can create a lot of burden for that team to keep up with over time. The second challenge we see is the sheer growth of log data overall. Last quarter, we surveyed over 270 folks who are responsible for their observability budgets. And we asked them how much data had grown over the past year. In a one-year time frame, they reported back data had grown on average 250%. Huge astronomical growth that we’re seeing basically across the board. It’s probably pretty obvious why this is bad. When you think about a log management or a SIEM tool, you’re often being charged based on the volume of data you capture in that tool.
As this data grows and grows, it becomes impossible to put all that data in your log management platform without breaking your budget. So the question becomes: "How do I get my team the data they need to support their use cases without completely breaking my budget?" And that’s a really hard question to answer.
That’s what we’re seeing across the board. And let’s talk a little bit about how we solve this as an organization. So Chronosphere’s Telemetry Pipeline is a vendor agnostic platform that collects your data from any source. It preprocesses that data in flight and it streams it to any destination. In essence, you can think of it as basically a control plane that gives you control over your data from the point of its creation all the way to its final destination.
Introducing Chronosphere’s Telemetry Pipeline
Riley Peronto: A couple things I want to just double click on here. First, I use the term "vendor agnostic." Chronosphere’s Telemetry Pipeline is a standalone product offering. You do not need to send your logs to the Chronosphere platform. We absolutely support that, but we also work with New Relic, Datadog, Elastic, whatever tools you have in your stack, we can stream logs to those outputs.
The second thing I want to spell out is that this platform is built on Fluent Bit. Our technical co-founder and Field Architect, Anurag Gupta, who is joining me in today’s webinar, is actually the creator of Fluent Bit. With that, the platform is built on open source and open standards. So there’s no new proprietary agents you need to deploy. There’s no vendor lock in of any sort.
It’s all open source and open standards. On the right hand side of the screen here, we list a few ways in which you can pre-process data with Chronosphere. That’s really the focal point of the webinar today. And it’s just the tip of the iceberg what I have on screen here. I do want to call out this reduced row though, so a common driver for adopting telemetry pipeline technology is to right size your data footprint to reduce logging costs. When customers approach us with this use case in mind, big and small customers have been able to save at least 30%, often 50%, or 60% on their logging costs, both through the reduction of cloud egress fees, as well as the reduction of indexing and retention in your log management platform.
So, we have a lot of really powerful outcomes when you’re using Chronosphere Telemetry Pipeline.
Now, before we dive into the actual demo, where we’re going to show you the different ways in which you can process data, I wanted to help you guys visualize how this technology might fit into your stack. This is actually an amalgam of a couple different customers that we work with. And what we did is, we looked for common infrastructure patterns just to give you a visualization of what this might look like before and after.
Streamlining observability migration
What we’re looking at here, the customer is streaming log data to a few different backends here. They’re leveraging proprietary vendor specific log collection agents. They’re pushing their Windows and Microsoft Exchange logs through Logstash. They have a Kafka pipeline they need to maintain. They also have syslog outputs from their network devices and their firewalls. So if you think about this, take a step back. That’s five to six different resources they need to manage. And their team needs to make sure that everything’s configured correctly and is pushing logs to the right destination.
It’s just a big burden for teams to keep up with all of this work and all of these different resources. When you look at the after, effectively what we’ve done is we’ve consolidated collection and routing. All within the chronosphere telemetry pipeline. You can manage this all from one place and we’ve made everyone’s lives a lot easier in doing so – at the same time we’re pushing compute processing upstream so that you can manipulate or shape or transform your data in different ways, whether to better support your use cases and drive specific outcomes or to support the different backends you support. The one thing I want to call out here, if you might notice that we swapped out one of the backends for Dynatrace, what Chronosphere can effectively do is it allows you to collect data from one resource or one agent and push that to a different backend.
This can streamline the migration from one backend to another. Another thing I want to call out is just the efficiency at which we execute these use cases. We recently benchmarked our pipeline against a competing telemetry pipeline in the market. And in doing so, we performed a sizing calculation and we found that Chronosphere uses about 20x less infrastructure resources than this other pipeline.
What this means for customers, is that you’re spending less money, you have a lower TCO on your telemetry pipeline technology, you’re spending less time managing infrastructure resources, and you’re getting better performance as you’re pushing large volumes of data downstream. Overall, we’re able to do this in a very high performance, efficient manner.
With that, Anurag, I’ll hand it off to you and we can talk a bit more about specific processing use cases and we can dive into the actual demo.
Processing data with Chronosphere’s Telemetry Pipeline
Anurag Gupta: Great to chat with everyone today and I’m pretty excited to show the demos for what we can do with processing both security and observability data. And it really comes down to use cases. When we as Chronosphere think about processing, we think of it in two ways. One is: “How do we go address these use cases that you see on the screen? But the second is also, how do you make it so that folks in the organization have rapid access?” To go test, try, and build processing that can adhere to basic things, advanced things, or log reduction use cases. So, I’ll throw a little bit about that and how we think of it and why that’s different from some of the other players in the market. The first bit is — what is basic processing.
Processing is generally an overloaded term, right? If you think of a computer processor it’s doing everything and anything under the sun. And so the way that we like to think about it is if I have a stream of data coming in, that data might need to be changed a little bit, have some additional fields added, some context to make it faster for debugging and easier to understand. We might want to do a little bit of parsing, and what that means is taking that log and, or taking that data stream and deriving it into different keys. We might do some, or a redaction, we might do some additional transformations on top of that really changing the data as it comes through.
Really your very basic, I have one message in, one message out type of use case. Now, advanced processing is where things start to get fun, and, if folks are in the security play market and you’ve heard about this OCSF, open comma schema format for security logs, or you have heard of other schemas like SAF, CEF, or other things that your SIEM requires, those schemas can be pretty heavy to, to go and enact, and typically what you’ll find is.
You have to use this proprietary method, this bespoke integration to get that schema in a way that your detections, your alerts are going to be as efficient as possible. The nice thing with what we have out of the basic processing, these out of the box templates is we give you some of that stuff.
You can easily test, try things like Windows events and see what those transformations look like to get it into those schemas. The second is enrichments from other data sources. We think of your basic process, it’s just one message in, one message out. Advanced processing can say one message will look up against an API.
Go retrieve a CSV file that has some indicator of compromise type scenario. And then use that as part of the data feed. And then last but not least, when we talk about log reduction: Log reduction can be made up of both these basic and advanced use cases. How do we do some small transformations that can be really impactful? I’ll show a couple of examples of that. And then two, with some of the advanced processing where you’re essentially doing some stream processing, like aggregation, dedupe, sampling. How can we summarize that to, really make an impact on some of the costs that we’ve started to incur. Those are the use cases for processing. I’m going to dive straight into the demo here.
Looking at the log data
Anurag Gupta: The nice thing about that is the pipeline only has one outbound connection to our management plane. It’s going, it’s checking if I’ve enacted a new processing rule. It goes, pulls that down and then restarts the pipeline to now adhere to that new processing piece.
Reducing log data
Anurag Gupta: And now, we have key values. So I have a parsed field that has a bunch of subkeys here, like your path, user, code and I can see how many of these error codes might be 404s, 500s et cetera. You can see my reduction also went down. It went from 89 percent to 74%, but now I have a lot of views and a lot of insights into what’s going on with the logs that are more meaningful than just the blob of text that exists.