CloudOpsGuide
observability

Grafana Loki: Log Aggregation Without the Elasticsearch Tax

Intermediate
10 minutes
October 2026
CloudOpsGuide Team

Grafana Loki: Log Aggregation Without the Elasticsearch Tax

Metrics tell you something is wrong; logs tell you what. Grafana Loki brings Prometheus's label-based philosophy to logs — cheap to run, native to Grafana, and designed for the Kubernetes era.

Table of Contents

Why Loki (and not Elasticsearch)

Elasticsearch indexes the full text of every log line — powerful search, heavy infrastructure. Loki indexes only labels (app, namespace, pod) and stores log content compressed in chunks. Result: a fraction of the storage cost and operational weight, at the price of less flexible full-text search. For "show me the errors from api pods in prod", Loki is the sweet spot.

Architecture in one breath

  • Promtail / Alloy — agent that tails logs and pushes them with labels
  • Loki — stores chunks, serves queries
  • Grafana — queries it with LogQL

Deploy on Kubernetes with Helm

helm repo add grafana https://grafana.github.io/helm-charts
helm install loki grafana/loki \
  -n monitoring --create-namespace \
  --set deploymentMode=SingleBinary \
  --set loki.commonConfig.replication_factor=1
helm install alloy grafana/alloy -n monitoring
# Alloy auto-discovers pod logs and ships them to Loki

Add Loki in Grafana as a data source (http://loki-gateway.monitoring:80) and Explore gives you live log tailing.

LogQL: the queries you'll actually run

# all logs from one app
{app="api", namespace="prod"}

# filter lines containing error
{app="api"} |= "error"

# exclude health-check noise
{app="api"} |= "error" != "healthcheck"

# parse JSON logs, filter on a field
{app="api"} | json | level="error"

# count errors per minute — logs become metrics
sum(count_over_time({app="api"} |= "error" [1m]))

That last pattern is the killer feature: turn log lines into time series, then alert on them without a separate pipeline.

Alerting from logs

# recording/alerting rule for Loki ruler
- alert: AppErrorSpike
  expr: |
    sum(count_over_time({app="api"} |= "ERROR" [5m])) > 50
  for: 5m

Retention and storage

limits_config:
  retention_period: 720h        # 30 days
compactor:
  retention_enabled: true

Point chunk storage at S3/GCS/MinIO for durability — local disk is for labs only.

Label discipline: the thing everyone gets wrong

Loki performance lives and dies by label cardinality. Good labels: app, namespace, env, pod. Terrible labels: request_id, user_id, trace_id — every unique value creates a new stream. Never put unbounded values in labels; extract them at query time with | json or logfmt instead.

Metrics for symptoms, logs for causes, traces for journeys — Loki covers the middle pillar without the Elasticsearch tax.

Related Articles


Last Updated: October 2026 Author: CloudOpsGuide Team Difficulty: Intermediate Estimated Reading Time: 10 minutes