Skip to content
Techsense Developers
TrustLet's Talk
Insights
Data, AI & MLOps6 min readOct 3, 2026

How to Monitor LLM Applications with Prometheus and Grafana: A Production Telemetry Guide

Effective LLM monitoring starts with treating your model calls like any other production dependency: instrument them, scrape the metrics, and visualize the results. The fastest path to…

Effective LLM monitoring starts with treating your model calls like any other production dependency: instrument them, scrape the metrics, and visualize the results. The fastest path to production-grade observability is to export four classes of telemetry from your application: request latency, token usage, error rates, and cost. You expose these as Prometheus metrics, scrape them on an interval, and render them in Grafana. That combination gives you real-time visibility into how your LLM features behave under load, where they fail, and how much they cost per request, without locking you into a proprietary vendor dashboard.

This guide walks through the full pipeline. I will cover what to measure, how to instrument a Python service, how to configure Prometheus to scrape it, and how to build Grafana dashboards and alerts that catch problems before your users do.

Why LLM Monitoring Differs From Standard APM

Traditional application performance monitoring assumes deterministic, low-variance operations. A database query either returns in a few milliseconds or it times out. LLM calls break that assumption in several ways, which is why generic APM is not enough on its own.

  • Latency is highly variable. Response time depends on prompt length, output length, model size, and provider-side queueing. A p50 of 800ms can sit alongside a p99 of 12 seconds.
  • Cost scales with tokens, not requests. Two calls to the same endpoint can differ in cost by an order of magnitude depending on context size.
  • Failures are semantic, not just structural. A 200 OK can still contain a refusal, a truncated completion, or a hallucinated answer. HTTP status codes do not capture this.
  • Rate limits and quotas are external. You are monitoring a dependency you do not control, so you need to track provider throttling explicitly.

Good LLMOps observability means capturing both the mechanical signals (latency, errors) and the economic and quality signals (token usage, truncation, retries). Prometheus handles the time-series side well, and Grafana gives you the panels and alerting on top.

The Core Metrics to Export

Before writing any code, decide what to measure. I recommend these as a baseline for Prometheus LLM telemetry:

  1. Request count by model, endpoint, and status. A counter.
  2. Request latency as a histogram so you can compute percentiles.
  3. Token usage split into prompt tokens and completion tokens. Counters.
  4. Estimated cost derived from token counts and your price table. A counter.
  5. Error count by error type (rate limit, timeout, server error, validation).
  6. Active in-flight requests. A gauge, useful for spotting saturation.

The Prometheus data model uses labels to slice these metrics. Keep label cardinality under control: model and status are safe, but never label by user_id or raw prompt text. That will blow up your time-series database.

Instrumenting a Python LLM Service

Here is a minimal but production-shaped example using the official prometheus_client library. Assume a FastAPI service that wraps an LLM provider.

from prometheus_client import Counter, Histogram, Gauge

# Request volume and outcome
LLM_REQUESTS = Counter(
    "llm_requests_total",
    "Total LLM requests",
    ["model", "status"],
)

# Latency distribution. Buckets tuned for LLM response times.
LLM_LATENCY = Histogram(
    "llm_request_duration_seconds",
    "LLM request latency in seconds",
    ["model"],
    buckets=(0.5, 1, 2, 4, 8, 16, 32, 64),
)

# Token accounting
LLM_PROMPT_TOKENS = Counter(
    "llm_prompt_tokens_total", "Prompt tokens consumed", ["model"]
)
LLM_COMPLETION_TOKENS = Counter(
    "llm_completion_tokens_total", "Completion tokens produced", ["model"]
)

# Estimated spend in USD
LLM_COST = Counter(
    "llm_cost_usd_total", "Estimated cost in USD", ["model"]
)

# Concurrency
LLM_INFLIGHT = Gauge(
    "llm_requests_inflight", "In-flight LLM requests", ["model"]
)

Now wrap the call site. The important detail is that you record metrics in a finally block so failures are counted, and you derive cost from the usage object the provider returns.

import time

# Price per 1K tokens. Keep this in config, not hardcoded.
PRICING = {
    "gpt-4o": {"prompt": 0.0025, "completion": 0.01},
}

async def call_llm(model: str, messages: list) -> dict:
    LLM_INFLIGHT.labels(model=model).inc()
    start = time.perf_counter()
    status = "error"
    try:
        response = await client.chat.completions.create(
            model=model, messages=messages
        )
        status = "success"

        usage = response.usage
        LLM_PROMPT_TOKENS.labels(model=model).inc(usage.prompt_tokens)
        LLM_COMPLETION_TOKENS.labels(model=model).inc(usage.completion_tokens)

        price = PRICING.get(model, {"prompt": 0, "completion": 0})
        cost = (
            usage.prompt_tokens / 1000 * price["prompt"]
            + usage.completion_tokens / 1000 * price["completion"]
        )
        LLM_COST.labels(model=model).inc(cost)

        return response
    except RateLimitError:
        status = "rate_limited"
        raise
    finally:
        LLM_LATENCY.labels(model=model).observe(time.perf_counter() - start)
        LLM_REQUESTS.labels(model=model, status=status).inc()
        LLM_INFLIGHT.labels(model=model).dec()

Finally, expose the metrics endpoint. With FastAPI you can mount the Prometheus ASGI app:

from prometheus_client import make_asgi_app

app.mount("/metrics", make_asgi_app())

Your service now publishes token usage metrics, latency, and cost at /metrics in the Prometheus exposition format.

Configuring Prometheus to Scrape

Prometheus pulls metrics on a schedule. Add a scrape job pointed at your service. A 15-second interval is a reasonable starting point for most LLM workloads.

# prometheus.yml
global:
  scrape_interval: 15s

scrape_configs:
  - job_name: "llm-service"
    metrics_path: /metrics
    static_configs:
      - targets: ["llm-service:8000"]
        labels:
          environment: production

If you run on Kubernetes, use the Prometheus Operator and a ServiceMonitor instead of static targets so scraping follows your pods automatically. Confirm the target is healthy under Status > Targets in the Prometheus UI before moving to Grafana.

Building Grafana Dashboards

With data flowing, the next step is Grafana dashboards that turn raw counters into operational insight. Add Prometheus as a data source, then build panels using PromQL.

Latency percentiles

The histogram_quantile function computes percentiles from your latency histogram:

histogram_quantile(
  0.95,
  sum(rate(llm_request_duration_seconds_bucket[5m])) by (le, model)
)

Plot p50, p95, and p99 on the same time-series panel, one query per quantile. Divergence between p50 and p99 tells you whether slowness is systemic or tail-driven.

Error rate

Express errors as a proportion of total traffic so the signal stays meaningful as volume changes:

sum(rate(llm_requests_total{status!="success"}[5m]))
/
sum(rate(llm_requests_total[5m]))

Cost burn rate

Track hourly spend per model to catch runaway loops or prompt bloat:

sum(rate(llm_cost_usd_total[1h])) by (model) * 3600

Token throughput

sum(rate(llm_completion_tokens_total[5m])) by (model)

Group these into a single dashboard with rows for Traffic, Latency, Errors, and Cost. Use Grafana variables for model and environment so one dashboard serves every deployment.

Alerting on What Matters

Dashboards are for investigation. Alerts are for response. Define alerting rules in Prometheus or Grafana for the conditions that signal real degradation.

  • Elevated error rate. Fire when the error ratio exceeds a threshold over a sustained window.
  • Latency SLO breach. Alert when p95 crosses your service-level target.
  • Cost spike. Alert when hourly burn exceeds budget, which often indicates a retry storm or a prompt regression.
  • Provider throttling. Alert on any sustained rate of status="rate_limited".

A sample Prometheus alerting rule:

groups:
  - name: llm-alerts
    rules:
      - alert: HighLLMErrorRate
        expr: |
          sum(rate(llm_requests_total{status!="success"}[5m]))
          / sum(rate(llm_requests_total[5m])) > 0.05
        for: 10m
        labels:
          severity: warning
        annotations:
          summary: "LLM error rate above 5% for 10 minutes"

The for clause prevents flapping by requiring the condition to hold before firing.

Where to Go From Here

This stack gives you the operational backbone. To extend it toward full quality monitoring, add instrumentation for truncation (when finish_reason is length), retry counts, and cache hit rates if you use a prompt or response cache. For teams standardizing across many services, OpenTelemetry can feed both traces and metrics into the same pipeline, which pairs well with the approach above.

If you are formalizing this across a platform, our Data, AI & MLOps capabilities cover the architecture patterns involved, and we have seen these observability requirements play out differently across regulated and high-volume industries, where cost governance and auditability carry extra weight.

The goal is simple: no LLM feature should reach production without the same telemetry discipline you apply to every other critical path.

FAQ

What is the difference between LLM monitoring and standard application monitoring?

Standard monitoring tracks latency, throughput, and errors for deterministic operations. LLM monitoring adds economic and semantic signals: token usage, per-request cost, truncation, retries, and provider rate limiting. LLM latency is also far more variable, so percentile tracking matters more than averages.

Why use Prometheus and Grafana instead of a managed LLM observability tool?

Prometheus and Grafana are open source, vendor-neutral, and likely already running in your infrastructure. You own the data, control retention, and avoid per-event pricing that scales poorly with high request volume. Managed tools can complement this stack, but the Prometheus pipeline keeps your core telemetry portable.

How do I avoid high cardinality in my LLM metrics?

Label only by low-cardinality dimensions such as model, status, and environment. Never label metrics by user ID, request ID, or prompt content. High-cardinality labels create millions of time series and can overwhelm your Prometheus server. Put per-request detail in logs or traces instead.

How can I track LLM cost accurately with Prometheus?

Derive cost from the token usage object the provider returns on each call, multiplied by a price table you keep in configuration. Increment a cost counter per request, then use PromQL rate() queries to visualize hourly or daily burn. Keep the price table versioned so you can audit historical spend.