ELK Stack: From the Classic Trio to a Modern Observability Platform

Modern distributed systems and cloud environments generate hundreds of millions of log entries every day. Turning scattered events into information that engineers can investigate on demand is a core challenge for operations and backend teams. ELK Stack, a name synonymous with centralized log analytics, has shaped how teams observe their systems for more than a decade.

But “ELK” today does not always mean the original three components. It may refer specifically to Elasticsearch, Logstash, and Kibana, or simply serve as a team’s familiar shorthand for its Elastic logging platform—even if Beats or Elastic Agent now handles log collection.

Distinguishing these uses makes it easier to understand how the classic architecture relates to later products, without assuming that removing Logstash means a system no longer counts as ELK.


What Is ELK Stack? The Roles of Its Three Core Components

ELK Stack is a data processing pipeline composed of three independent tools: Elasticsearch, Logstash, and Kibana. It brings together raw events from individual hosts and organizes them into clearly defined fields that engineers can query and chart.

Each component has a distinct role:

Put simply, Logstash prepares the data, Elasticsearch stores and searches it, and Kibana presents the results.

Is All of ELK Open Source?

All three tools began as open-source projects, and Elastic now leads their development and distribution. Being an “open-source project” does not mean there is no company behind it—a company can employ core developers, set the product direction, and offer both free and paid features.

Their current licenses are not all the same. Logstash remains under Apache 2.0. Portions of the Elasticsearch and Kibana source code are available under a choice of AGPLv3, SSPL 1.0, or Elastic License 2.0, while the default official distributions use Elastic License 2.0.

AGPLv3 is an open-source license approved by the Open Source Initiative (OSI), whereas SSPL and Elastic License 2.0 are not OSI-approved—even though both still publish the source code. That is why access to source code does not necessarily mean the software uses an open-source license.

Calling ELK “three open-source tools” is therefore imprecise. A more accurate description is that all three have roots in the open-source ecosystem and are developed under Elastic’s leadership, but the licensing terms vary across components and official distributions. The “Licensing Changes and the Evolving Ecosystem” section below covers the 2021 licensing shift and why it led to OpenSearch.


The Classic Data Flow: Investigating Problems with Nginx Logs

To understand how the three components work together, consider web access logs in the classic architecture. Text logs from a single host are easy to read, but analyzing activity across dozens of servers requires centralized, structured data.

 [ Sources: Nginx / App / Syslog ]
               | Raw events
               v
 [ Logstash: Input -> Filter -> Output ]
   - Input: file, syslog, beats
   - Filter: grok, mutate, date
   - Output: elasticsearch
               | Structured documents
               v
 [ Elasticsearch: Cluster and storage ]
   - Inverted indexes and columnar storage
   - Distributed search and aggregations
               | Query results
               v
 [ Kibana: Visualization and analysis ]
   - Discover / Dashboard / Alerts

From Alert to Diagnosis: A Practical Workflow

Suppose a production system raises an alert about a rising HTTP 500 error rate. In a traditional environment without centralized logs, the on-call engineer must log in to each host and run tail or grep. If a container restarts, the logs needed for the investigation may disappear altogether.

In the classic ELK workflow, engineers can locate the problem by querying logs:

  1. Parse log fields: Logstash receives log text and uses Grok patterns to turn unstructured strings into fields such as status, request_path, and duration. The host field and other host metadata may be added by a collector such as Beats.
  2. Filter to the alert window: The engineer sets the time range in Kibana Discover to match the alert window. Elasticsearch filters the indexed data by @timestamp to retrieve the relevant events.
  3. Group by path and host: Filter for status >= 500, then group by request_path and host in real time to quickly discover that all the errors come from the /api/checkout endpoint on a particular host.

ELK’s core value is that engineers can query logs spread across hosts and written in different formats through a single interface, then correlate them to find the problem.

How Does Grok Turn a Log Line into Fields?

Grok is a Logstash filter that extracts fields from text. Think of it as a collection of named regular-expression building blocks: built-in patterns such as IP, NUMBER, and HTTPDATE each recognize a common format. Engineers combine these patterns instead of writing a long, hard-to-read regular expression from scratch.

Grok does not infer what each piece of text means. Engineers still need to specify which part of the log format represents an IP address, timestamp, request path, or status code, and name the extracted values. Suppose Nginx produces the following access log, where the final number is the request processing time in seconds:

 203.0.113.7 [22/Sep/2026:14:32:10 +0800] "POST /api/checkout HTTP/1.1" 500 842 0.731

The corresponding Grok pattern could be:

 %{IP:client_ip} \[%{HTTPDATE:timestamp}\] "%{WORD:http_method} %{URIPATHPARAM:request_path} HTTP/%{NUMBER:http_version}" %{NUMBER:status:int} %{NUMBER:bytes:int} %{NUMBER:duration:float}

%{IP:client_ip} matches text using the built-in IP pattern and stores the result in client_ip. %{NUMBER:status:int} recognizes a number, names it status, and converts it to an integer. Logstash uses this pattern to produce a structured event:

{
  "client_ip": "203.0.113.7",
  "http_method": "POST",
  "request_path": "/api/checkout",
  "status": 500,
  "bytes": 842,
  "duration": 0.731
}

Next, the date filter can convert the original timestamp into the standard @timestamp field. If a log line does not match the Grok pattern, Logstash tags it _grokparsefailure by default, allowing engineers to review those parse failures separately.

This process does not make Logstash understand the meaning of a log. It extracts and names pieces of text according to a known format. A line that was previously just readable text becomes a set of fields that can be filtered, grouped, and aggregated.


How ELK Evolved: From Independent Projects to the Elastic Platform

ELK was not a software suite planned by a single team. It took shape gradually, as three open-source projects that each solved a different problem converged.

 Around 2010
 [Elasticsearch] -+
 [Logstash     ] -+--> [ELK Stack]  (2013, one company)
 [Kibana       ] -+        |
                           v
                  [Elastic Stack]  (2016, 5.0 / Beats)
                           |
                           v
                  [License change]  (2021, OpenSearch fork)
                           |
                           v
                  [AGPLv3 added]  (2024)

Bringing Three Independent Projects Together

Around 2010, the three components emerged as independent open-source projects. Elasticsearch was a distributed search engine built on Apache Lucene, exposing Lucene’s search capabilities as a service accessible through RESTful APIs and JSON. Logstash established the “Input ➔ Filter ➔ Output” pipeline model. Kibana provided a visual query interface specifically for Elasticsearch.

In 2013, the main maintainers of all three projects joined the same company. The ELK Stack architecture took shape and quickly became an industry standard for centralized log analytics.

From Beats and Elastic Stack to Elastic Agent

A common early approach was to run Logstash on every server, continuously read new entries from local log files, parse them, and forward them to Elasticsearch. The problem was that every host needed a full Logstash JVM process running continuously. Even when used only to ship logs, it consumed memory and CPU.

Elastic subsequently introduced Beats, lightweight collectors dedicated to specific data types, such as Filebeat and Metricbeat. In 2016, Elastic aligned the component versions at 5.0 and officially named the suite Elastic Stack.

As host counts grew, managing individual Beats configuration files became labor-intensive. In recent years, Elastic has introduced Elastic Agent to combine multiple collection capabilities in a single agent. Paired with Fleet in Kibana, it lets teams distribute configurations and manage upgrades centrally. With collection now handled by a single agent, Logstash shifted to a backend role, focusing on consolidating data from different sources and handling complex extract, transform, and load (ETL) workflows.

The platform’s scope has gradually expanded beyond logs. Elastic Agent can collect logs, metrics, and traces, while Kibana has grown from its early role as a visualization interface into a shared entry point for exploring data, configuring alerts, and managing observability and security features.

Elastic Stack has evolved beyond simply adding tools to ELK. What began as a fixed, three-stage logging pipeline is now a collection of tools that teams can combine according to their data sources and processing needs.

Licensing Changes and the Evolving Ecosystem

In 2021, Elastic replaced Apache 2.0 licensing with a dual-license model using SSPL and the Elastic License. AWS responded by forking the codebase to create OpenSearch. In 2024, Elastic announced AGPLv3 as an additional open-source option for its source code.


Benefits and Costs: What to Consider Before Adoption

ELK offers flexible search and analytics, but it requires real investment in hardware and operations.

Beyond Centralized Logs: What Can the Same Pipeline Do?

Centralized logging remains ELK’s most straightforward use case. Once events from web servers, applications, operating systems, and network devices are organized into common fields, engineering teams can search them along a shared timeline, compare behavior before and after deployments, track capacity trends, and run audit queries.

Security teams can use the same approach for failed logins, permission changes, firewall events, and endpoint events. Time-series data such as product clickstreams and transaction events can also flow through the same pipeline.

The modern Elastic platform integrates metrics and traces into the same query interface. After receiving a latency alert, engineers can see which services a request passed through and how much time it spent in each, then use the request ID to look up the application logs.

This does not mean all data must pass through Logstash or use exactly the same indexing strategy. It means logs, metrics, and traces can be correlated through common fields and time ranges. Moving from viewing logs centrally to investigating problems across signals is a key step in ELK’s evolution into an observability platform.

Core Strengths

Operational Costs to Expect

The costs do not end when the cluster starts running. Teams must decide which fields to index, how long to retain data, when to move it to cheaper storage tiers, and how to handle backups, upgrades, permissions, and failure recovery. Managed services can offload some host management and upgrade work, but they do not eliminate the costs driven by data volume, indexing strategy, and retention policies.

These storage and operational costs helped drive the emergence of lighter logging solutions suited to cloud-native environments.


Is ELK Outdated? Its Role in Modern Architectures

To assess whether ELK still fits, separate two questions: whether its old deployment patterns remain appropriate, and whether the problems it solves still exist.

The heavyweight approach of running Logstash on every host and fully indexing all logs is no longer the default for most teams. But the core needs—collection, normalization, retention, search, and visualization—have not changed. Today’s observability systems offer more ways to route that data.

Modern Data Flows Do Not Always Need Logstash

Classic ELK places Logstash at the sole entry point, while modern architectures choose different paths based on processing complexity. Data with a stable format that needs only basic field transformations can go directly from Elastic Agent, Beats, or an application to Elasticsearch. Elasticsearch’s built-in ingest pipelines can then parse or transform fields before indexing.

Logstash takes on the central processing role when teams need to consolidate data from multiple sources, perform complex conditional transformations, or send data to multiple destinations.

The OpenTelemetry Collector offers another entry point: applications can use a common standard to generate and collect logs, metrics, and traces, then send that data to Elastic or other backends.

OpenTelemetry addresses how telemetry data is generated, collected, and exported. Elasticsearch, OpenSearch, and Loki address how it is stored and queried. They operate at different layers; teams do not have to choose between them.

 [ Data sources ]
       |
       v
 [ Collection: Elastic Agent / Beats / OTel Collector ]
       |
       v
 [ Processing (as needed): Logstash / ingest pipeline ]
       |
       v
 [ Storage and query backends:
   Elasticsearch / OpenSearch / Loki ]
       |
       v
 [ Exploration and visualization:
   Kibana / OpenSearch Dashboards / Grafana ]

A real architecture does not need every layer, and not every combination of tools is possible. Teams choose paths based on their collection methods, transformation needs, and backend support.

Comparing Roles in the Modern Observability Ecosystem

SolutionCore mechanism and roleUse cases and trade-offs
Elasticsearch / Elastic StackInverted indexes plus on-disk columnar structures for aggregations; a full-featured observability platform covering logs, metrics, and tracesSupports arbitrary full-text searches and multidimensional real-time aggregations; higher storage and Java cluster operating costs
OpenSearchAn open-source fork of Elasticsearch 7.10.2 that retains its distributed architectural foundationsSuits teams that need Apache 2.0 license compatibility or deep integration with AWS managed services
Grafana LokiIndexes only labels and stores compressed log content in object storage such as S3Suits lower-cost, long-term retention of cloud-native container logs; queries narrow the scope by labels before scanning log content line by line, making it less suited to complex full-text exploration
OpenTelemetry (OTel)Industry-standard telemetry specifications and tools focused on generating, collecting, and exporting dataUnifies the collection layer and connects to backends such as Elastic and Loki, reducing vendor lock-in

Start with Query Patterns, Then Choose a Backend

ELK’s strength is that engineers can begin with a full-text search for an error message even when they do not yet know where the problem lies, then narrow the results by service, version, host, and status code. If incident investigations frequently require this kind of exploratory querying, and the data comes from systems with varied formats, Elasticsearch’s full-text indexes and multidimensional aggregations offer clear value.

Conversely, if data formats are consistent, most queries revolve around a few known labels, and the main goal is to retain large volumes of Kubernetes logs at a lower cost, fully indexing everything may not be economical. A design like Loki’s, combining label indexes with object storage, will often fit those needs better.

Before choosing a solution, answer a few questions: How much do data formats vary across sources? Do you need arbitrary full-text search? Can common queries narrow the scope using labels first? How long must the data be retained? Can your team standardize field formats and manage the index lifecycle—from creation through tiered storage and deletion—as well as backups and upgrades? These answers do more to determine the right architecture than which tool is newer.


Conclusion: Choose Based on Query Needs and Operational Capacity

Today, “ELK” may refer to the classic trio that took shape in 2013 or to the modern Elastic platform. Its separation of collection, storage, and querying remains fundamental to understanding centralized log analytics.

When choosing a modern architecture, weigh the following principles:

Understanding how classic ELK divides collection, processing, and querying—and comparing that with what each component can do today—helps teams assemble an observability architecture that fits their needs.