Service Intelligence
Service Intelligence turns raw telemetry into a service-level view of your estate. You group the entities and log sources that make up a service, define its health as a set of LPQL-based KPIs, map how services depend on one another, and let LogPulse watch each KPI for anomalies. It runs on the same LPQL and ClickHouse engine as search and the SIEM, so a service's health signals are just queries you can click into and pivot from.
Overview
A service is a logical thing you operate, a checkout API, an authentication provider, a payments worker, assembled from the hosts, containers, and log sources that produce its telemetry. Service Intelligence gives each service a single health status rolled up from its KPIs, and collects them into one hub so you can see, at a glance, what is healthy and what needs attention.
It lives in the Observability hub, alongside Health Checks, Entities, and Maintenance windows. The hub lists every service with its current status and summary counts; opening a service drills into its KPIs, members, dependencies, anomalies, and settings.
Service-level health
One rolled-up status per service, derived from its KPIs, so you watch services rather than scattered metrics.
KPIs from LPQL
Any LPQL search becomes a health signal with thresholds, no separate metrics pipeline to feed.
Dependency-aware
Map upstream and downstream services to see how a failure propagates across the estate.
Anomaly-watched
Each KPI is baselined and watched for deviations, so drift is caught before it breaches a threshold.
Defining a Service
Create a service, give it a name and description, then tell LogPulse which telemetry belongs to it. Membership is resolved automatically and stays current as entities come and go, so you define the service once rather than maintaining a static list.
Membership & Scope
A service is scoped one of two ways:
| Scope | How membership works |
|---|---|
| Entity labels | Match entities by their labels (for example team, app, or env). Membership rules select every entity whose labels match, and the set updates automatically as entities are discovered. |
| Log source | Scope by one or more log-source names. The service is defined by the sources it owns and has no entity members. |
KPIs
A KPI (Key Performance Indicator) turns an LPQL search into a health signal for the service. The query produces a value, a chosen field is read as the metric, and thresholds map that value to a severity. KPIs are the building blocks of a service's health.
Each KPI defines the search, the value field to read, a unit, and its thresholds. The service's scope is put in front of the search when it runs, so the search itself only describes the measurement. For example, a 5xx error-rate KPI:
# KPI: 5xx error rate (%)
status=*
| eval is_error=if(status>=500, 1, 0)
| stats count as total, sum(is_error) as errors
| eval error_rate=if(total>0, round(errors/total*100, 2), 0)Read error_rate as the value field, set the unit to %, the threshold direction to above, and a warning at 1 with a critical at 5. The KPI is evaluated on a schedule; its latest value and severity are shown on the service, and you can chart it over 1h / 6h / 24h / 7d windows.
The KPI Editor
New KPI on a service, or Edit in a KPI card's menu, opens the editor for that one KPI. The same menu disables or enables a KPI, runs it once without storing anything (Run now), and deletes it after a confirmation.
| Field | Meaning |
|---|---|
| Name / Display name | The machine name (unique within the service) and the name shown on KPI cards and in the mobile widgets. |
| Description | One or two sentences on what the KPI measures. Shown on the KPI card. |
| Unit | A short label, at most 16 characters: ms, %, req/min, errors. Shown with every value: "%" directly after the number, other units after a space. |
| Decimals | 0 to 6 digits after the decimal point. Empty shows whole numbers as they are and fractions with at most two digits. Values from 10 000 are shortened (12.5K). |
| Search | The LPQL that produces the number, usually ending in stats. |
| Value field | Which column of the result is the metric. Empty uses the count. |
| Split field | Optional. One value per entity (pod, endpoint, queue); needs a matching "by" clause. Thresholds and baselines then apply per entity. |
| Time window | How far back each run looks: -5m, -1h, -7d. Pick one from the list or type a custom -<number><m|h|d>. |
| Run interval | How often the KPI runs, at least every minute. |
Templates
A new KPI can start from the template library: the four golden signals (latency, traffic, errors, saturation) per kind of service. Picking a template fills the search, value field, unit, decimals, window, interval and suggested thresholds. Everything stays editable, and the thresholds are starting points to tune against your own traffic in the threshold preview.
| Kind of service | Templates |
|---|---|
| Web / API | Latency (p95), request rate, 5xx error rate, rejected requests (429/503), server errors per endpoint |
| Database | Query latency (p95), query rate, errors, slow queries, refused connections |
| Queue / worker | Processing time (p95), throughput, failed jobs, queue depth per queue |
| Batch job | Run duration, successful runs and failed runs over 24 hours |
| Authentication | Login latency (p95), successful logins, login failure rate, account lockouts |
| Kubernetes workload | Log rate, error logs per container, crash and OOM signals, pods logging |
| Generic | Log rate, errors, error rate, hosts reporting |
status and duration (milliseconds) from the Web Access data model, or job_status for queues and batch jobs. If your logs name them differently, change the search or map the fields with a data model.Advanced Settings
| Setting | Meaning |
|---|---|
| Enabled | A disabled KPI is not measured and does not count towards the service health. |
| Weight | 0 to 100, default 1. Stored as the relative importance of the KPI; the service health is currently the most severe enabled KPI, so the weight does not change it yet. |
| Anomaly sensitivity | 0.25x to 4x, default 1x. Higher flags smaller deviations from the baseline, lower only flags large ones. |
| Baseline days | 1 to 14, default 5. Days of history a KPI needs before anomalies are raised; until then it shows as learning. |
| Baseline lookback | 7 to 56 days, default 28. How much history the expected range per weekday and hour is learned from. Longer is steadier, shorter follows a changed service sooner. |
| Anomaly cooldown | 5 to 1440 minutes, default 30. Quiet time before the same anomaly is raised again. |
| AI investigation | Off by default. When on, an anomaly of at least the chosen severity (low, medium, high, critical) starts an AI investigation. |
Thresholds & Severity
A threshold has a direction and up to two breach levels. The direction decides whether high or low values are bad; the levels decide how bad.
| Field | Meaning |
|---|---|
| Direction | "above" flags values over the threshold (error rate, latency); "below" flags values under it (success rate, throughput). |
| Warning | The value at which the KPI turns Warning, an early signal, not yet an incident. |
| Critical | The value at which the KPI turns Critical. The service is breaching its objective. |
| Value field | Which field from the LPQL result is read as the metric. |
Service Health
A service's overall status is rolled up from its KPIs: the most severe KPI wins, so a single Critical KPI makes the whole service Critical. The Observability hub summarizes the estate with counts of Critical, Warning, and Healthy services, and you can filter the list by status to triage the ones that matter.
| Status | Meaning |
|---|---|
| Healthy | All KPIs are within their thresholds. |
| Warning | At least one KPI has crossed its warning level. |
| Critical | At least one KPI has crossed its critical level. |
| Unknown | No KPIs have reported yet (newly created, or awaiting first evaluation). |
SLOs & Error Budgets
Health says how a service is doing right now. A service level objective (SLO) says how reliable it has to be over time: "99.9% over 30 days". The 0.1% that may go wrong is the error budget. The SLOs tab of a service shows, per SLO, the SLI now (the measured percentage), how much of the budget is left, how fast it is being spent, and a burn-down of the budget over the window. SLOs are evaluated every 5 minutes.
Two Ways to Measure
| Kind | What is counted | Use it for |
|---|---|---|
| Time a KPI is good | Every run of a KPI is one period. A period is good when the KPI is not critical (or, if you choose so, only when it is normal). Runs that failed are left out. | Availability or latency you already track as a KPI; works back as far as the KPI has results (90 days). |
| Good events / all events | Two LPQL searches, scoped to the service and counted every 5 minutes: the good events and all events. | Request-based objectives such as "99.5% of requests succeed". Counting starts when the SLO is created. |
For an events SLO, write the searches as filters and let LogPulse count them. Compare numbers with | where, because a comparison in the search expression compares text:
Good events: status=* | where status < 500
All events: status=*A search that ends in its own | stats ... as count is used as written. The editor previews what the SLI would have been over the last 7 days before you save. Targets run from 50% to 99.999%; 100% is refused because it leaves no budget to manage. Windows are rolling: 7, 28, 30 or 90 days.
Budget & Status
| Status | Meaning |
|---|---|
| Within budget | More than 10% of the error budget is left and nothing is burning. |
| At risk | Less than 10% of the budget is left, or a burn-rate alert condition holds. |
| Breached | The budget is spent. The card shows by how much it is overspent. |
| No data | Nothing was measured in the window yet. |
Time a service spends under a maintenance window counts on neither side, so planned work does not burn budget. A breached or at-risk SLO is listed among the reasons on the service's health panel, but it does not change the health score: health is about now, an SLO looks back over weeks.
Burn-Rate Alerts
The burn rate is how fast the budget is spent compared with spending it evenly over the window: at 1 the budget lasts exactly the window, at 14.4 a 30-day budget is gone in about two days. LogPulse uses the multi-window rule from the Google SRE workbook, so an alert fires quickly and stops soon after the problem does:
| Alert | Condition |
|---|---|
| Burning fast | Burn rate above 14.4 over the last hour and over the last 5 minutes. |
| Burning | Burn rate above 6 over the last 6 hours and over the last 30 minutes. |
| Burn stopped | Neither condition holds any more. |
Alerts go to the notification channels of the service (email, webhook, Slack, web push, mobile push): a fast burn counts as critical and a slow burn as warning for a channel's minimum severity. When the service is set to open incidents, a fast burn opens one and it resolves when the burn stops. A new alert waits 30 minutes after the previous one, a slow burn that turns fast is reported right away, and nothing is sent while the service is under maintenance. Burn-rate alerts can be switched off per SLO.
Dependencies
Services rarely fail in isolation. You can record upstream and downstream relationships between services and view them as a dependency graph, so when a service degrades you can see what it relies on and what relies on it. This turns a single red KPI into context: is checkout unhealthy because checkout broke, or because the payments service it depends on did?
KPI Anomaly Detection
Thresholds catch values you can name in advance; anomaly detection catches the ones you cannot. Each KPI is given a baseline, a per-KPI statistical profile of its normal range that accounts for daily and weekly seasonality, and LogPulse flags when the current value departs from it, even while it is still inside the static thresholds.
Anomalies surface on the service's Anomalies tab with the affected entity, the time the deviation started, and a severity. A per-KPI sensitivity setting (in the KPI editor's advanced settings) controls how far from baseline a value must drift before it is flagged. Raise it to catch subtle drift, lower it to only flag dramatic swings.
The baseline is learned per weekday and per hour, on the clock of your organization: set the baseline time zone in the security settings, under Anomaly scoring. With the default, UTC, a rhythm that follows local office hours shifts an hour when the clocks change.
Feedback That Tunes
The thumbs on an anomaly are not a survey. Not an issue lowers the sensitivity of that KPI by one step (to 85%, never below 0.25x); real issue raises a KPI that feedback lowered back up, never past the value it had before. The anomaly says what happened: "Sensitivity adjusted to 0.85 after your feedback". Setting the sensitivity by hand in the KPI editor always wins.
A KPI that is called not an issue three times at the same weekday and hour (the nightly batch, the Monday-morning login wave) stops raising anomalies for that hour. The rule appears under learned patterns, where you can switch it off. Static thresholds are never suppressed.
View logs on an anomaly, and in the menu of a KPI card, opens the search the KPI runs, service scope included. For a split KPI the anomaly's link is narrowed to the entity that deviated.
Change Events
"What changed just before this?" is the first question about any deviation. Change events answer it: deploys, config changes and feature-flag flips show as vertical markers on the KPI charts of a service, in a list on its Changes tab, and on an anomaly as "Changes in the 2 hours before". KPI cards also draw the previous period as a faint dashed line and say how the average compares ("+12% vs previous period").
Add a change by hand on the Changes tab, or let your CI pipeline post one after every deploy. Create a personal access token with the services:propose scope and call:
curl -X POST https://api.logpulse.io/api/v1/changes \
-H "Authorization: Bearer $LOGPULSE_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"kind": "deploy",
"service": "checkout-api",
"title": "Deploy checkout-api",
"version": "'"$GIT_SHA"'",
"url": "'"$CI_PIPELINE_URL"'"
}'| Field | Meaning |
|---|---|
| title | Required. What changed, in a few words. |
| kind | deploy (default), config, feature_flag, incident or other. |
| service / serviceId | The service by name (any casing) or by id. Leave both out for a change that concerns the whole organization; it then shows on every service. |
| version | A version, tag or commit shown next to the title. |
| url | An http(s) link to the pipeline run, the pull request or the change ticket. |
| occurredAt | ISO 8601 with offset; defaults to now. |
| description | Optional longer text, shown on the Changes tab. |
A service name that does not exist answers 404 instead of silently marking every chart. Read changes back with GET /api/v1/changes?serviceId=…&from=…&to=… (a session, or a token with services:read).
Entity 360
Every member of a service is an entity with its own unified view. Entity 360 combines an observability lens, the services, KPIs, and telemetry tracked for the entity, with a security lens of accumulated risk and reputation. The same view is reachable from the Observability hub and from the security side, so an SRE and an analyst arrive at the same full-context page from different directions. See the Security Monitoring (SIEM) page for the security lens.