Executive Summary
Log management is a cost-control discipline as much as a technical one — index everything at premium rates and the bill, not the data, becomes the problem you can’t ignore.
Splunk, Elastic, Datadog, and Grafana Loki cover centralized log aggregation, search, and retention from very different cost and architecture angles — from Splunk’s powerful, premium-priced analytics to open-source-rooted Elastic, SaaS observability suites, and Loki’s deliberately lean label-indexing model. They share the same job, but the dividing line is economics at scale: log volumes grow relentlessly, and ingestion, indexing, and retention pricing is what ultimately shapes the decision.
This guide provides a vendor-neutral evaluation framework for 8 leading platforms, weighing cost model at real log volumes, search and analytics power, and fit within broader observability so you can control spend and retention rather than absorb the bill shock that defines this category.
Why Log Management & Analysis Matters for Enterprise Strategy
Log-management selection is dominated by cost at scale: volumes grow relentlessly and pricing tied to ingestion, indexing, and retention can outrun the value, so a platform’s economics and tiering options weigh as heavily as its search power. The strategic questions are whether logs belong in a standalone tool or a broader observability platform, and self-managed versus SaaS — each trading operational burden against predictable cost.
Cheaper indexing models, tiered and archival storage, and the consolidation of logs with metrics and traces into unified observability are reshaping how teams buy log management. Weigh each platform on cost predictability and data tiering at your volume and on AI-assisted analysis, because the long-term expense lives in ingestion and retention, not in the initial license.
Architecture & Sourcing Decision
Log management is almost never a true build-vs-buy question — running your own ELK or Loki cluster at scale is buying the open-source engine and then paying for it in engineering, not avoiding the cost. The real decisions are architectural: standalone log tool versus part of an observability suite versus part of the SIEM; index-heavy (search everything fast) versus index-free or label-only (cheap storage, slower ad-hoc search); and whether a telemetry pipeline should sit in front of whatever you pick to govern volume before it ever hits a priced backend. Frame the choice around where your logs are consumed — the SOC, the SRE on-call, or the auditor — and what you can afford to retain hot.
Because ingest and retention are the cost engine of this category, the highest-leverage decision is often not the store at all but whether you route through a pipeline first. The four camps below rarely compete head-to-head; most shortlists end up choosing across them.
| Your Situation | Recommended Path | Rationale |
|---|---|---|
| Logs feed security detection as the primary use case | SIEM-led platform (Splunk, Microsoft Sentinel) | When the buyer is the SOC, correlation rules, threat content, and analyst workflow matter more than raw log search; a SIEM-anchored platform avoids a second copy of the data and a second tool to operate. |
| You already run metrics and traces in one suite | Logs inside the observability platform (Datadog, Dynatrace) | Pivoting from a trace to its logs in one click shortens MTTR more than any standalone log UI; the trade is per-GB pricing and lock-in, so pair it with index-free tiers for high-volume, low-value logs. |
| Spend is exploding at a premium-priced backend | Add a telemetry pipeline (Cribl, OpenTelemetry Collector) | A pipeline trims, samples, and routes telemetry before it hits a priced store — the fastest lever on ingest cost without ripping out the incumbent. Remember it governs data; it is not itself the log store. |
| Kubernetes-native, cost-sensitive, search is mostly recent | Index-free / label-only store (Grafana Loki) | If most queries are “show me this pod’s logs for the last hour,” label-only indexing slashes storage cost; accept slower full-text search across cold, unstructured data as the trade-off. |
| Long retention & full-text forensics across a mixed estate | Index-heavy with tiered storage (Elastic, Sumo Logic) | Deep, fast search over months of logs needs a real index; control the bill with hot/warm/frozen tiering and searchable snapshots rather than keeping everything on hot storage. |
Key Capabilities & Evaluation Criteria
Weight these domains against your own log volume, retention obligations, and who actually consumes the data. For most enterprises the cost-and-tiering model now outranks raw search features that older RFPs over-index on — because at scale the platform you can afford to keep logs in matters more than the one with the cleverest query language. Score ingest economics and search power together; a platform that is brilliant at search but unaffordable to feed is not a winner.
| Capability Domain | Weight | What to Evaluate |
|---|---|---|
| Cost & Tiering Model at Volume | 25% | Unit of pricing (per-GB ingest, compute/SVC, credits, host, consumption) and how the bill behaves as volume grows; hot/warm/cold/frozen or archive tiers, index-free or low-cost retention tiers, rehydration cost and latency, and ingest controls or sampling to cap spend |
| Search, Query & Analytics Power | 20% | Full-text vs. label-only vs. schema-on-read; query language depth (SPL, ES|QL/KQL, LogQL, DQL, SQL) and learning curve; performance on cold and high-cardinality data; aggregations, pattern detection, and live-tail for real-time troubleshooting |
| Pipeline, Ingestion & OpenTelemetry | 20% | Agent footprint and collector strategy; native OpenTelemetry/OTLP ingest; parsing, enrichment, masking, routing, and reduction before storage; ability to fan data out to multiple destinations and to swap backends without re-instrumenting sources |
| Security, Compliance & Retention | 15% | RBAC/SSO, field-level masking of sensitive data, immutable/WORM audit retention, and tamper-evident logging; SOC 2 / ISO 27001 / FedRAMP coverage and data residency; SIEM or threat-detection content if logs feed security |
| Observability Correlation & Architecture | 10% | Whether logs sit standalone, inside an observability suite (correlated with metrics and traces), or inside the SIEM; trace-to-log pivots, unified context, and how cleanly it slots beside tooling you already run |
| Operations, Scale & AI Assist | 10% | Self-managed vs. SaaS operational burden, multi-tenant and multi-cluster scale, ingest spikes and back-pressure handling; API/IaC coverage; and AI-assisted features such as anomaly detection, log clustering/patterns, and natural-language query |
Vendor Landscape
The market no longer divides cleanly into “log tools.” Logs now sit at the intersection of observability and security, and platforms approach them from one of three directions: SIEM-led players where logs feed threat detection first; observability suites where logs are one pillar beside metrics and traces; and storage-and-search engines optimized for keeping and querying logs cheaply at scale. Cutting across all three is a distinct layer — the telemetry pipeline — that governs volume before it reaches any priced backend. Most shortlists end up comparing across these camps rather than within one.
Ownership has shifted markedly: Cisco completed its roughly $28B acquisition of Splunk in March 2024, folding the category’s most established platform into a networking-and-security giant; Sumo Logic was taken private by Francisco Partners in 2023; and New Relic was taken private by Francisco Partners and TPG the same year. The through-line is consolidation around platforms, with independent pure-plays increasingly the exception. We profile the eight platforms that most often make enterprise shortlists; New Relic remains a credible observability-suite option for teams already in its ecosystem, and developer-centric tools such as Better Stack serve smaller engineering teams well below this tier.
Strengths: The most mature enterprise log platform, with the deepest query language (SPL), a vast app and integration ecosystem, and the strongest security analytics and SIEM story in the category. Cisco ownership (deal closed March 2024) adds network telemetry and a path to tie logs to the broader Cisco security and observability portfolio. Workload (SVC) pricing offers an alternative to pure per-GB ingest for estates with lots of rarely-searched data. Considerations: Sits at the premium end on cost, and licensing is genuinely complex — ingest vs. workload pricing is a decision in itself. Deployment and tuning reward certified expertise, and migrating a large on-prem estate to Splunk Cloud is a real project. Buyers should track how Cisco continues to integrate and package the portfolio.
Strengths: Cloud-native SIEM with log management built on Azure Monitor / Log Analytics, now extended by a data lake tier (GA in 2025) and auxiliary/basic log plans that let high-volume, low-value security logs land cheaply and graduate to the analytics tier when needed. Deep, native integration with Microsoft 365, Entra, and Defender makes it the path of least resistance for Microsoft-centric security teams, with strong KQL analytics and built-in detection content. Considerations: Strongest when logs serve security; it is not a general-purpose APM/observability log tool. Cost spans Sentinel and the underlying Log Analytics meters, so model both; tiering and table-plan management (analytics vs. basic/auxiliary vs. data lake) adds operational nuance. Value concentrates inside the Azure and Microsoft security estate.
Strengths: The reference index-heavy engine for fast, flexible full-text search and analytics over logs, with a strong open-source heritage — Elasticsearch and Kibana returned to an OSI-approved open-source license (AGPL) alongside Elastic License v2 in 2024. Mature data tiering (hot/warm/cold/frozen) with searchable snapshots keeps long retention affordable, ES|QL simplifies querying, and the same platform spans logs, security, and search use cases. Considerations: Self-managing Elasticsearch at scale demands real operational expertise — cluster sizing, sharding, and upgrades are not trivial, and “free” OSS becomes an engineering line item. Elastic Cloud is competitive but grows with data; some security and ML features sit in paid tiers. The 2021 license history left lingering caution and the AWS-backed OpenSearch fork as an alternative.
Strengths: Logs fully correlated with metrics, traces, and the rest of the Datadog platform, so on-call engineers pivot from a failing trace straight to its logs. Logging without Limits decouples ingest from indexing — ingest everything, index selectively, archive the rest — and Flex Logs adds a low-cost tier for high-volume retention and historical search. Polished pipelines, parsing, and UX. Considerations: Per-GB ingest plus indexing means costs can climb fast and unpredictably without disciplined ingest controls, and the breadth of the platform makes lock-in real. Rehydrating archived logs for investigation carries its own scan-and-index cost. Best economics require actively managing which logs are indexed vs. flexed vs. archived.
Strengths: Not a log store — a telemetry pipeline that sits between sources and destinations to collect, reduce, enrich, mask, and route log and event data, with broad out-of-the-box integrations. Its core value is governing volume and cost before data hits a priced backend like Splunk, Sentinel, or Datadog, and freeing buyers to send the same data to multiple destinations or swap stores without re-instrumenting sources. Considerations: Frame it correctly: Cribl does not retain or search logs as a system of record, so it complements rather than replaces a store — it is another component (and cost) in the chain, and you still need a destination. It adds an operational layer to design and run, and overlaps with the OpenTelemetry Collector, which covers some of the same routing/processing ground at no license cost.
Strengths: Label-based, index-free design — it indexes only metadata labels, not log content — for substantially lower storage cost than full-text engines. Kubernetes-native, with LogQL familiar to PromQL users and tight integration with Grafana dashboards alongside Prometheus metrics and Tempo traces. Recent engine work targets larger analytical and high-cardinality queries, and Grafana Alloy provides an OpenTelemetry-native collection path. Considerations: The index-free trade-off shows on broad, ad-hoc full-text searches over older data, which are slower than on an index-heavy engine; queries lean on getting your labels right. Self-managing Loki at scale is non-trivial, and Grafana Cloud pricing introduces its own model. Less of an out-of-the-box enterprise SIEM/compliance story than Splunk or Elastic.
Strengths: Born-in-the-cloud, multi-tenant SaaS spanning log analytics and Cloud SIEM with no infrastructure to run. Its Flex pricing model removes separate ingest and indexing charges and meters by analytics tier (continuous, frequent, infrequent), giving a credit-based lever to match cost to how often data is actually queried. Strong for cloud and DevSecOps teams that want operations and security logging on one managed platform. Considerations: As a SaaS-only platform it suits cloud-first estates more than heavy on-prem footprints needing local low-latency access. Credit-based Flex pricing is flexible but takes modeling to forecast, and tier choice (frequent vs. infrequent) directly shapes both cost and query latency. Now privately held under Francisco Partners (since 2023); weigh roadmap continuity as you would with any take-private.
Strengths: Log management built on Grail, a schema-on-read data lakehouse that ingests logs without up-front indexing and queries them with DQL, unified with Dynatrace metrics, traces, and topology and enriched by its Davis AI for automated anomaly detection and root-cause. A “Retain with Included Queries” option lets teams hold data at a fixed cost with bounded querying for shorter retention windows, aiding predictability. Considerations: Greatest value comes when you adopt the wider Dynatrace platform; as a standalone log tool it is less of a natural fit, and DQL plus the consumption (DPS) model is a learning curve. Pricing has several consumption dimensions (ingest, retention, query) to model. Typically a premium, platform-level commitment rather than a tactical log purchase.
Pricing Models & Cost Structure
The headline rate matters far less than the unit of measure — per-GB ingested, compute (Splunk SVCs), credits by analytics tier, host/agent, or platform consumption — because that unit determines what you pay as volume grows, and log volume always grows. The hidden multipliers live in indexing, retention tier, and the cost (and latency) of querying or rehydrating cold data after the fact. Model every shortlisted platform against your projected volume twelve months out, separating cheap-to-ingest from expensive-to-index-and-keep, and price in any pipeline layer you add to control it.
| Vendor | Pricing Model | Relative Tier | Key Cost Drivers |
|---|---|---|---|
| Splunk (Cisco) | Per-GB ingest or workload (SVC) compute; term/subscription | Premium | Daily ingest volume or SVC compute, search/retention footprint, premium apps (ES, ITSI), on-prem vs. Splunk Cloud |
| Microsoft Sentinel | Consumption: analytics vs. basic/auxiliary vs. data-lake tier | Moderate | GB ingested per tier, Log Analytics retention, table-plan choice, commitment tiers, queries against lake/auxiliary data |
| Elastic | Resource/consumption (Elastic Cloud) or self-managed subscription | Moderate | Data volume and hot/warm/cold/frozen tier mix, cluster compute and storage, edition (security/ML in paid tiers), self-run FTE |
| Datadog Log Management | Per-GB ingest + per-million-events indexing; Flex Logs tier | Premium | Ingest volume, indexed events and retention, Flex Logs storage, rehydration scans, broader platform (metrics/traces) coupling |
| Cribl | Volume-based (data processed) subscription | Moderate | Throughput of data through the pipeline, worker footprint, products in use (Stream/Edge/Lake); offsets downstream store cost |
| Grafana Loki | Open-source (self-run) or Grafana Cloud usage-based | Lower | Object storage volume, self-managed compute and FTE, or Grafana Cloud ingested GB and retention; low index overhead by design |
| Sumo Logic | Credit-based Flex; metered by analytics tier | Moderate | Credits consumed by tier (continuous/frequent/infrequent), ingest volume, retention, Cloud SIEM bundle, edition |
| Dynatrace | Platform consumption (DPS): ingest + retention + query | Premium | GiB ingested and retained on Grail, query (scanned-GiB) volume or Retain-with-Included-Queries option, broader platform adoption |
Implementation & Migration
Sequence the rollout by data value and cost exposure, not by what is easiest to onboard. Get the noisiest, highest-volume sources under a pipeline and tiering policy early — that is where the bill is made — and prove the critical search and detection use cases before you fan out to every source.
Inventory log sources, volumes, and growth, and tag each by use (security detection, troubleshooting, audit/compliance). Run the POC on real ingest, model the 12-month bill per shortlisted platform, and decide your indexing-vs-archive and retention policy up front with security and finance at the table.
Deploy collection — ideally OpenTelemetry-native — and, where volume warrants, a telemetry pipeline to parse, mask sensitive fields, reduce, and route before storage. Configure hot/warm/cold or index-free tiers, wire identity (SSO/RBAC), and onboard the first high-value sources rather than everything at once.
Codify the queries, dashboards, alerts, and (if security-led) detection content that justify the platform. Validate query and forensic performance on cold/archived data, set retention and WORM/audit holds to compliance requirements, and confirm trace-to-log or SIEM correlation works for the real on-call and analyst workflows.
Extend to remaining sources, then establish standing cost governance: ingest controls and budgets, periodic review of what is indexed vs. archived, and tier rebalancing as volume shifts. Decommission the legacy tool only after parity is proven, and review actual spend against the original model.
Selection Checklist & RFP Questions
Use this checklist during evaluation to verify the capabilities that actually decide cost and usefulness at scale — not just a generic feature tick-list.