Prometheus Recording Rules: Precompute Expensive Queries

Khimananda Oli 9 min read DevOps
Prometheus Recording Rules: Precompute Expensive Queries

By Khimananda Oli | Last reviewed: September 2026

Slow Grafana panels and timing out API calls are often symptoms of unoptimized PromQL, not insufficient hardware. Prometheus recording rules: precompute expensive queries at scrape time so dashboards and alerts read simple, pre-aggregated vectors instead of scanning millions of raw samples on every request. This guide shows you exactly how to identify candidates, write correct rules, and validate them before they reach production. If you are building your observability stack from scratch, start with the Prometheus and Grafana full monitoring stack to ensure your foundation supports rule evaluation efficiently.

Raw Metrics (TSDB)High Cardinalityhttp_requests_total{instance, path, method}Recording Rule EngineEvaluates every 15ssum(rate(http_requests_total[5m]))by (service)Precomputed ResultDashboard / AlertFast Vector Lookupjob:http_requests:rate5m< 5ms Latency
Prometheus recording rules precompute expensive queries by materializing aggregations into new time series for fast downstream consumption.

How do Prometheus recording rules precompute expensive queries?

Recording rules function as a scheduled ETL layer inside the Prometheus server itself. Unlike standard queries that execute only when a user loads a dashboard or an alert fires, recording rules run continuously at a defined interval. The engine evaluates the PromQL expression against the current TSDB state and writes the resulting vector back into the database as if it were a scraped metric.

This mechanism fundamentally changes the performance profile of your monitoring stack. When you define meaningful SLIs and SLOs, you often need aggregations across thousands of pods or instances. Without recording rules, every dashboard refresh forces Prometheus to decompress and process all matching raw samples. With recording rules, the heavy lifting happens once per evaluation cycle. The dashboard then queries a low-cardinality, pre-aggregated metric that returns almost instantly. For teams managing compliance or audit-ready infrastructure, this predictability is essential; you cannot have critical SLO dashboards timing out during an incident review.

The evaluation lifecycle

  1. Trigger: The rule group timer fires based on the configured interval (e.g., every 15 seconds).
  2. Fetch: Prometheus reads the necessary raw samples from the TSDB for the lookback window specified in the query.
  3. Compute: The PromQL engine executes the aggregation, rate calculation, or join.
  4. Store: The resulting sample is appended to the TSDB under the new metric name defined in the record field.
  5. Expose: Subsequent queries for that metric name read directly from storage, bypassing the original complex computation.

When should you create recording rules instead of direct queries?

A common mistake is creating recording rules for everything "just in case." This bloats your TSDB and increases memory pressure. You should only apply Prometheus recording rules: precompute expensive queries when specific performance or usability thresholds are crossed. In my experience auditing monitoring stacks for SOC 2 compliance, I look for three concrete signals that justify a new rule.

Signal 1: Query latency exceeds 500ms

If a panel in Grafana consistently takes more than half a second to render, the underlying query is too heavy for interactive use. Use the Prometheus UI's "Query History" or check the prometheus_http_request_duration_seconds histogram to identify slow endpoints. If the p95 latency for /api/v1/query_range is creeping up, it is time to offload that computation.

Signal 2: High cardinality aggregations in alerts

Alerting rules must evaluate quickly to avoid missing firing windows. If your alert condition involves a sum(rate(...)) over a high-cardinality metric like http_requests_total with many label combinations, move that aggregation to a recording rule. Alerts should reference simple thresholds against recorded metrics, not perform math on raw data during critical moments. See alerting with Prometheus Alertmanager for structuring reliable alert pipelines.

Signal 3: Repeated sub-expressions

If multiple dashboards or alerts share the same complex base query—such as a specific error rate calculation filtered by region and environment—extract it. Duplicating complex PromQL leads to drift and maintenance nightmares. A single recorded metric becomes the source of truth for that business logic.

Complex PromQL IdentifiedIs query latency > 500ms ORused in multiple dashboards?NoYesKeep as Direct QueryCreate Recording RuleValidation Checklist✓ Naming convention followed✓ Labels preserved correctly✓ Interval matches scrape✓ No circular dependencies
Decision framework for determining when Prometheus recording rules precompute expensive queries versus leaving them as direct queries.

How do you configure and name Prometheus recording rules correctly?

Configuration lives in YAML files referenced by the rule_files directive in prometheus.yml. While syntax is straightforward, naming and structure determine long-term maintainability. In 2026, the community-standard naming convention remains the most effective way to prevent confusion between raw and computed metrics.

Naming convention: level:metric:operations

Adopt the hierarchical naming scheme recommended by Prometheus documentation. This makes autocomplete useful and prevents collisions.

  • Level: The aggregation scope (job, instance, region, or cluster).
  • Metric: The base metric name being aggregated.
  • Operations: The functions applied (rate5m, sum, avg, count).
groups:
- name: api_performance_rules
  interval: 15s
  rules:
  # GOOD: Clear hierarchy and operation
  - record: job:http_requests_total:rate5m
    expr: sum(rate(http_requests_total[5m])) by (job)

  # GOOD: Preserves critical labels for SLO filtering
  - record: job:http_errors:ratio_rate5m
    expr: |
      sum(rate(http_requests_total{code=~"5.."}[5m])) by (job)
      /
      sum(rate(http_requests_total[5m])) by (job)

  # BAD: Ambiguous name, loses context
  - record: error_rate
    expr: sum(rate(http_requests_total{code=~"5.."}[5m])) / sum(rate(http_requests_total[5m]))

Preserving labels for future aggregation

Always retain the labels you might need to filter or group by later. If you aggregate away the region label in a global recording rule, you can never drill down by region without reverting to the raw query. When defining SLIs, keep dimensions like service, endpoint, or environment unless cardinality explicitly forbids it. Refer to defining meaningful SLIs and SLOs for guidance on which dimensions matter for reliability tracking.

What are the trade-offs between recording rules and direct PromQL queries?

Understanding the cost model is vital for capacity planning. Recording rules are not free; they exchange query-time CPU for storage space and ingestion-time CPU. The table below summarizes the operational trade-offs I consider when designing monitoring architectures for multi-cloud environments.

FactorDirect PromQL QueryRecording Rule
Read LatencyHigh (scales with data volume)Low (constant time lookup)
Write AmplificationNoneIncreases TSDB size (~10-30%)
CPU ProfileSpiky (user-driven)Constant (background evaluation)
Data FreshnessReal-time (current scrape)Delayed by evaluation interval
Historical AccuracyRecalculated on schema changeFrozen at evaluation time
Cardinality ImpactTemporary during queryPermanent stored series

The "Historical Accuracy" row deserves emphasis. If you fix a bug in your PromQL or change a label matcher, existing recorded data remains unchanged. Direct queries always reflect the current logic against raw history. For compliance-critical metrics where audit trails must be reproducible, document your rule versions alongside your infrastructure code.

How do you validate and debug recording rules before deployment?

Never deploy untested recording rules to production. A malformed rule can silently produce incorrect data for weeks until someone notices the dashboard looks "off." Follow this validation workflow to ensure correctness.

Step 1: Syntax and unit testing

Use promtool check rules to validate YAML syntax and basic semantic correctness. More importantly, write unit tests using promtool test rules. Define input samples, expected output samples, and assert that your rule produces the exact values you expect. This catches edge cases like division by zero or empty vector matches before they affect users.

# example_test.yml
evaluation_interval: 15s
tests:
  - interval: 15s
    input_series:
      - series: 'http_requests_total{job="api", code="200"}'
        values: '0+10x10'
      - series: 'http_requests_total{job="api", code="500"}'
        values: '0+1x10'
    promql_expr_test:
      - expr: job:http_errors:ratio_rate5m
        eval_time: 2m
        exp_samples:
          - labels: '{job="api"}'
            value: 0.0909090909

Step 2: Ad-hoc verification in Prometheus UI

Before adding the rule to your config, run the expr directly in the Prometheus Graph tab. Compare the result graph against your mental model. Check that labels appear as expected. Verify that the range vector covers enough samples for rate() to be meaningful—at least 4 data points within the bracket duration.

Step 3: Monitor rule evaluation health

After deployment, watch prometheus_rule_evaluation_failures_total and prometheus_rule_group_last_duration_seconds. If evaluation duration approaches your interval, your rules are too heavy and will start lagging. Also monitor prometheus_tsdb_head_series to detect unexpected cardinality explosions from overly broad recording rules. Integrating these checks into your Prometheus metrics monitoring fundamentals ensures the system watches itself.

Before: Direct QueriesDashboard Load: 2.4s avgCPU Spikes: 85% on refreshTSDB Reads: 2.1M samples/qAlert Eval: 1.8s (risk of delay)User Experience: FrustratingAfter: Recording RulesDashboard Load: 45ms avgCPU Steady: 12% backgroundTSDB Reads: 450 samples/qAlert Eval: 15ms (reliable)Storage Cost: +18% seriesTrade-off: Storage for Speed
Performance comparison demonstrating how Prometheus recording rules precompute expensive queries to reduce latency at the cost of additional storage.

Optimizing Your Monitoring Stack with Precomputation

Implementing Prometheus recording rules: precompute expensive queries is one of the highest-ROI optimizations available to platform engineers. It transforms monitoring from a fragile, reactive tool into a reliable, scalable foundation for decision-making. Start by profiling your slowest queries, apply the naming conventions strictly, and validate every rule with unit tests before merging. Remember that rules are code: version them, review them, and retire them when they no longer serve a purpose. If your team needs help designing a compliant, high-performance observability architecture or auditing an existing stack, reach out to discuss your monitoring strategy.

Frequently Asked Questions

They precompute expensive PromQL expressions and save results as new time series. This reduces query latency for dashboards and alerts by shifting calculation load from read time to write time during ingestion.

Direct queries on high-cardinality data cause timeouts and high CPU usage. Recording rules calculate results once per interval, making dashboard loads instant and reducing strain on the Prometheus storage engine during peak traffic.

Add a groups block to your rules file with rule_type set to record. Specify the new metric name under record and the PromQL expression under expr. Reload configuration using promtool check rules before applying changes to production.

Match the evaluation interval to your scrape interval, typically fifteen or thirty seconds. Setting it lower than the scrape rate creates gaps, while setting it significantly higher causes stale data and misleading trends in downstream dashboards.

Yes. By precomputing complex aggregations, you reduce CPU and memory pressure on Prometheus servers at query time. This often allows downsizing instance types or avoiding vertical scaling, directly lowering monthly compute bills for monitoring stacks.

Follow the level:metric:operations format. The level indicates aggregation scope like job or instance, metric is the base name, and operations describe transformations. This standard prevents naming collisions and makes rule purpose immediately clear to other engineers.

They can if labels are not dropped during aggregation. Always use without or by clauses to remove high-cardinality labels like pod_id or request_id. Uncontrolled label retention multiplies stored series and accelerates storage growth unnecessarily.

Use promtool test rules with unit test files defining input samples and expected outputs. Run these tests in CI pipelines to catch logic errors or regressions before rules reach production and corrupt historical metric data.

Prometheus logs the error and skips writing the sample for that interval. The metric shows gaps rather than incorrect values. Configure alerting on prometheus_rule_evaluation_failures_total to detect syntax errors or resource exhaustion issues immediately.

Yes, use promtool tsdb create-blocks-from to generate historical blocks from raw data. This populates past values for newly created rules, ensuring dashboards display continuous trends instead of starting empty from the deployment timestamp onward.

Recording rules generate new time series for future querying, while alerting rules evaluate conditions to fire notifications. Both share syntax but serve different purposes. Recording rules optimize performance; alerting rules trigger operational responses based on current state.

No. Separate them into distinct groups with appropriate evaluation intervals. Recording rules often need faster evaluation cycles than alerts. Mixing them forces unnecessary recalculations and complicates debugging when one rule type experiences performance degradation or failures.

Track prometheus_rule_group_last_duration_seconds and prometheus_rule_evaluation_failures_total. High duration indicates expensive expressions needing optimization. Failures suggest resource constraints or syntax issues. Set SLOs on evaluation latency to ensure rules complete within their scheduled interval.

Yes. Both systems support native Prometheus recording rule formats. Rules execute at the query layer or ingest path depending on architecture. Verify compatibility with your specific version, as some advanced PromQL functions may have limited support in distributed backends.

Avoid them for ad-hoc exploration, low-frequency queries, or metrics with extremely high churn rates. The storage overhead of persistent precomputed series outweighs benefits when queries run rarely or underlying data changes too rapidly for fixed intervals.