Grafana Loki: Log Aggregation Without the Elasticsearch Tax
Grafana Loki: Log Aggregation Without the Elasticsearch Tax
Metrics tell you something is wrong; logs tell you what. Grafana Loki brings Prometheus's label-based philosophy to logs — cheap to run, native to Grafana, and designed for the Kubernetes era.
Table of Contents
- Why Loki (and not Elasticsearch)
- Architecture in one breath
- Deploy on Kubernetes with Helm
- LogQL: the queries you'll actually run
- Alerting from logs
- Retention and storage
- Label discipline: the thing everyone gets wrong
Why Loki (and not Elasticsearch)
Elasticsearch indexes the full text of every log line — powerful search, heavy infrastructure. Loki indexes only labels (app, namespace, pod) and stores log content compressed in chunks. Result: a fraction of the storage cost and operational weight, at the price of less flexible full-text search. For "show me the errors from api pods in prod", Loki is the sweet spot.
Architecture in one breath
- Promtail / Alloy — agent that tails logs and pushes them with labels
- Loki — stores chunks, serves queries
- Grafana — queries it with LogQL
Deploy on Kubernetes with Helm
helm repo add grafana https://grafana.github.io/helm-charts
helm install loki grafana/loki \
-n monitoring --create-namespace \
--set deploymentMode=SingleBinary \
--set loki.commonConfig.replication_factor=1
helm install alloy grafana/alloy -n monitoring
# Alloy auto-discovers pod logs and ships them to Loki
Add Loki in Grafana as a data source (http://loki-gateway.monitoring:80) and Explore gives you live log tailing.
LogQL: the queries you'll actually run
# all logs from one app
{app="api", namespace="prod"}
# filter lines containing error
{app="api"} |= "error"
# exclude health-check noise
{app="api"} |= "error" != "healthcheck"
# parse JSON logs, filter on a field
{app="api"} | json | level="error"
# count errors per minute — logs become metrics
sum(count_over_time({app="api"} |= "error" [1m]))
That last pattern is the killer feature: turn log lines into time series, then alert on them without a separate pipeline.
Alerting from logs
# recording/alerting rule for Loki ruler
- alert: AppErrorSpike
expr: |
sum(count_over_time({app="api"} |= "ERROR" [5m])) > 50
for: 5m
Retention and storage
limits_config:
retention_period: 720h # 30 days
compactor:
retention_enabled: true
Point chunk storage at S3/GCS/MinIO for durability — local disk is for labs only.
Label discipline: the thing everyone gets wrong
Loki performance lives and dies by label cardinality. Good labels: app, namespace, env, pod. Terrible labels: request_id, user_id, trace_id — every unique value creates a new stream. Never put unbounded values in labels; extract them at query time with | json or logfmt instead.
Metrics for symptoms, logs for causes, traces for journeys — Loki covers the middle pillar without the Elasticsearch tax.
Related Articles
Last Updated: October 2026 Author: CloudOpsGuide Team Difficulty: Intermediate Estimated Reading Time: 10 minutes