IntegrationsAlerts and integrations

The rules that decide when Latent alerts, and how to send alerts to Slack, PagerDuty, email or your own webhook.

Pro and above

Latent checks your alert rules every minute. When a rule fires, it tells every destination attached to that rule, and when the condition clears, it tells them again. Destinations are Slack, PagerDuty, email and webhooks.

How alerts work

An alert rule reads one measure of your traffic over a trailing window and compares it with a threshold. A rule fires when three things hold:

  • It is turned on.
  • Its window holds at least its minimum number of answers.
  • The reading crosses the threshold.

It resolves as soon as any of the three stops holding. A quiet spell that drops the window under its minimum also resolves a firing rule, and when traffic returns and the reading still crosses, the rule fires again as a new episode. Each fire and each resolve goes to every enabled destination on the rule once, with no repeats while the rule stays firing.

Rules and destinations are set per project and apply to every model in it. A rule can be scoped to one model, and then it reads only that model's answers and, for the flag rate rules, uses that model's budget.

You manage both on the console's Policy page, in the Alert rules card. You can also connect a destination from Settings > Project > Integrations. Everyone on the project sees each rule's live reading against its threshold. Only admins can add, edit or turn rules on and off, and viewers see each destination without its address. Changes collect as a draft until you press Save, and a mistake is marked on the row it belongs to. If someone else saved first, nothing is saved: the console shows their version and who saved it, and you make your changes again and press Save.

What a rule can watch

Measure What it reads Threshold when left empty
Flag rate (escalation_rate) The share of scored answers flagged in the window, whether or not they go to the review queue Your flag budget
Flag rate under budget (escalation_rate_under) The same share, firing when it falls below a fraction of the budget, which catches flagging that has gone quiet Half the budget
Low confidence (low_confidence_rate) The share of answers Latent marks as low confidence 5%. To use another limit, set the rule's threshold
No answers (no_events) Seconds since the newest answer arrived. A new project has no reading until its first answer, so it never alerts before traffic starts The window, in seconds
Reviewed precision (precision_7d) Of the flagged answers your reviewers decided in the window, the share they confirmed as wrong 50%
Scoring time (scoring_host_ms_p95) The 95th percentile of meta.host_ms on your answers, in milliseconds. Hosted scoring and the sidecar's /score record it. The sidecar's streaming read and the in-band vLLM plugin do not, so on those paths the rule has no reading None. You set it
Review backlog (review_backlog) Answers sent for review in the window against your reviewers' daily capacity, scaled to the window Your reviewer capacity. The rule stays inactive until you set one
Drift (drift_ks) How far the risk scores of your latest answers have moved from your own first answers on the current calibration Worked out from the window and baseline sizes: 0.134 for the defaults

Reviewed precision and flag rate under budget fire only below their threshold. No answers, scoring time, review backlog and drift fire only above it.

On hosted plans the checks keep running after a Latent restart or redeploy, whether or not your app is sending answers or anyone has the console open.

The drift rule compares the last 200 scored answers with a baseline: the first 1,000 the model scored on its current calibration, extended by up to 2,000 more if that baseline fails its own consistency check. It fires after two readings in a row cross the threshold, each on new answers, and a rule left unscoped watches each model separately, so each model gets its own episode. When it fires, check the flag rate and precision. On hosted plans, request a calibration run from the console. When Latent runs in your environment, requantile, which resets the threshold on recent answers, or recalibrate.

The rules every project starts with

A new project starts with six rules and no destinations, so nothing is sent until you add a destination and attach it to a rule.

Rule Fires when At the start
Escalation over budget The flag rate is above your budget over 3 hours, with at least 20 answers Off
Escalation under half the budget The flag rate is under half your budget over 6 hours, with at least 50 answers Off
No events for an hour No answer has arrived for more than an hour Off
Low confidence over the drift limit The low-confidence share is over its limit over 24 hours, with at least 20 answers Off
Flagged for review over reviewer capacity More answers go to review in 24 hours than your reviewers can handle On, once you set reviewer capacity
Risk distribution drift from the deployment's baseline The drift reading crosses its threshold twice in a row On, once the model has 1,000 scored answers on its current calibration

A project holds up to 50 rules and 20 destinations. A rule can post to up to 20 destinations.

Slack

Pro and above

  1. In Slack, create an incoming webhook for the channel that should receive alerts, and copy its URL. It starts with https://hooks.slack.com/services/.
  2. In the console, open Policy and find the Alert rules card. Press Add destination, enter an id such as slack-alerts, set Kind to slack, paste the URL into Target, and press Apply. You can also connect it from Settings > Project > Integrations, which saves it at once.
  3. On each rule that should post to it, press Edit, tick the destination under Destinations, and press Apply. Then press Save on the card.

Each alert arrives as one line of text:

[latent] FIRED: Escalation over budget: escalation_rate 31.0% > 10.0% over 3 h (n=412)

The resolve reads RESOLVED in place of FIRED. A message about one model ends with model=<id>. A resolve sent because a rule was removed says so at the end.

PagerDuty

Pro and above

  1. In PagerDuty, add an Events API v2 integration to the service that should be paged, and copy its integration key.
  2. In the console, on the Alert rules card, press Add destination, enter an id, set Kind to pagerduty, paste the key into Target, and press Apply. A key is 16 to 128 letters, digits, _ or -.
  3. On each rule that should page, press Edit, tick the destination under Destinations, and press Apply. Then press Save on the card.

A rule that fires triggers an incident with severity warning and source latent-audit. When the rule clears, Latent resolves the same incident. Both events carry the dedup key latent-<rule id>-<model, or all>-<time the rule fired>, so one episode on one model is one incident. The incident's summary is the alert's one line of text. Its custom details hold the rule, the reading, the threshold, the number of answers read, and whether the rule fired or resolved.

Note Hosted plans post to PagerDuty's US endpoint. For an EU account, contact us. When Latent runs in your environment, set LATENT_PAGERDUTY_EVENTS_URL=https://events.eu.pagerduty.com/v2/enqueue on the service for an EU account.

Email

Pro and above

  1. In the console, on the Alert rules card, press Add destination, enter an id, set Kind to email, enter 1 to 20 addresses separated by commas in Target, and press Apply.
  2. On each rule that should email, press Edit, tick the destination under Destinations, and press Apply. Then press Save on the card.

The subject is the alert's one line of text, and the body is the alert as JSON, the same document a webhook receives.

Note When Latent runs in your environment, email needs a mail relay. Set LATENT_SMTP_URL on the service, as smtp://user:pass@host:port/?from=alerts@example.com, adding &starttls=1 for STARTTLS, or as smtps://... for TLS. Until it is set, the console offers email as unavailable and the service refuses to save an email destination.

Webhook

Pro and above

  1. In the console, on the Alert rules card, press Add destination, enter an id, set Kind to webhook, enter your endpoint's https:// or http:// URL in Target, and press Apply.
  2. On each rule that should call it, press Edit, tick the destination under Destinations, and press Apply. Then press Save on the card.

Latent POSTs the alert as JSON with Content-Type: application/json and User-Agent: latent-audit, and waits up to 5 seconds for a response. Any 2xx response counts as delivered. A redirect is refused, since following one would send the alert to whatever host the redirect names. The request carries no signature, so give the endpoint a URL that is hard to guess, and treat what it receives as a notice to check the console.

The host must be public. A Slack or webhook URL whose host is or resolves to a loopback, link-local, private, multicast, unspecified or reserved address, or whose host does not resolve, is refused when you save it, and the host is checked again before every POST, since DNS answers change.

Note When Latent runs in your environment and your relay is on your own network, set LATENT_ALERT_ALLOW_PRIVATE_WEBHOOK=1 on the service to allow private hosts. The URL must still be http or https and resolve, and redirects are still refused.

The payload

{
  "type": "latent.alert",
  "state": "fired",
  "service": "latent-audit",
  "ts": 1758400000.0,
  "rule": {"id": "budget_over", "name": "Escalation over budget", "metric": "escalation_rate", "op": ">",
           "value": null, "window_h": 3.0, "min_events": 20, "model": null,
           "baseline_n": null, "consecutive_ticks": null},
  "value": 0.31,
  "threshold": 0.10,
  "n_events": 412,
  "since": 1758400000.0,
  "message": "[latent] FIRED: Escalation over budget: escalation_rate 31.0% > 10.0% over 3 h (n=412)",
  "policy_version": 7
}
Field Meaning
type Always latent.alert
state fired or resolved
service Always latent-audit
ts When this notice was sent, in Unix seconds
rule The rule as configured: its id, name, measure, comparison, threshold (null for the measure's default), window in hours, minimum answers and model scope, plus the drift rule's baseline size and readings in a row
value The reading that crossed, or that stopped crossing
threshold The threshold in force, with defaults worked out
n_events How many answers the reading covers
since When this episode fired. The fire and its resolve carry the same value, so you can pair them
message The one line of text Slack and email show
policy_version The policy version in force when the notice was sent
detail Drift rules only: the statistics behind the reading, including the window and baseline sizes and how many readings in a row have crossed

Delivery and retries

  • A destination that fails is retried each following minute, up to 10 attempts for each fire and each resolve, counting the first. Retries of a fire stop once the rule resolves.
  • If the service restarts in the middle of a delivery and cannot tell whether it landed, it waits 10 minutes, then sends that notice again to the rule's destinations under the same fire time. PagerDuty keeps one incident, since the dedup key is the same. Slack, email and webhooks may receive the notice twice.
  • A destination that is turned off is recorded as disabled for that episode, and nothing is sent to it. Turning one off while a rule fires means it never gets the resolve, so an open PagerDuty incident stays open until you close it.
  • Turning a rule off while it fires resolves it at the next check, within a minute. Removing a firing rule sends its resolve to the destinations that still exist, so an open PagerDuty incident closes.
  • You cannot remove a destination while a rule still uses it.

The Alert history on the Policy page shows the newest 100 fires and resolves, with the reading and each destination's result: delivered, failed or disabled.

Acknowledging the budget monitor

The Policy page also carries the budget monitor, which is separate from the alert rules and off until an admin turns on its Alert rule switch. When it is on and your flag rate stays above budget for its set number of hours, 3 by default, Latent marks the calibration for revalidation and shows a banner on every console page.

An admin clears the mark with Acknowledge, on the banner or on the Monitor card under Advanced, where a note on the cause can be added. The card's Monitor history keeps who acknowledged it and any note. The monitor fires again on the next run of hours over budget.

Alert rules need no acknowledgement. Each one resolves by itself when its reading stops crossing.

Was this page helpful?

Updated 3 October 2026