Pings, endpoints, tokens
Reachability, status codes, response times and credential expiry — every target checked on a fixed cadence, every minute, from the moment you add it.
- ICMP and TCP targets
- HTTP checks with assertions
- Token expiry countdowns
Orion polls every service you run — pings, API endpoints, tokens, logs, metrics and clusters — and puts the state of all of it on one screen. When something breaks, the board shows it before the first support ticket lands.
Fleet verdict
1 service down
Uptime
99.97%
Mean RTT
50.3ms
Checks run
1,284,960
Fleet latency · last 20 min
peak 63.0 ms
Live stream
Signal wall · 32 checks
one cell per monitored thing
Reads the systems you already run
Most teams watch availability in one tool, logs in a second, metrics in a third and clusters in a fourth — then reconcile them by hand at two in the morning. Orion is the one screen those four were supposed to be.
Reachability, status codes, response times and credential expiry — every target checked on a fixed cadence, every minute, from the moment you add it.
Prometheus for metrics, Loki for logs, queried side by side instead of in three consoles with three logins and three mental models.
Kubernetes workloads, node vitals, NGINX gateways and WSO2 traffic, each with its own view and all rolling up into the same fleet verdict.
Point Orion at a failing window and it writes the root cause from the log lines themselves, finds the incident it resembles, and turns plain English into a query.
60 s
check cadence
every monitor, every minute
5
signal types
ping · API · token · log · metric
1
screen
the whole estate, one verdict
100%
self-hosted
your data never leaves your network
Dashboards are easy to make pretty and hard to make useful. Every surface in Orion is designed for the moment something is wrong and someone is trying to find out what.
Every check runs on its own schedule and writes its result. Orion keeps the pass and fail counts, so a monitor that is technically up but failing one check in twenty is impossible to hide.
Endpoint reliability · 24 h
60 s cadenceWhen latency moves, the reason is usually in the logs from the same two minutes. Orion queries Loki and Prometheus from one screen so you stop copying timestamps between tabs.
{app="checkout", env="prod"} |= "503"Orion drafts a written root cause from the log lines around a failure, cites the ones it used, and tells you which past incident this one resembles. Ask it a question in English and it writes the query.
Runs against a self-hosted model — your logs stay inside your network.
What happened. Every request to partners.orion.internal/v1/orders has returned 503 since 14:02. Three consecutive checks failed.
Why. The upstream began refusing connections two minutes after checkout-7d9f was OOMKilled and restarted — the connection pool was not re-established.
Orion runs where your services run. There is no agent to roll out and nothing to install on the things being watched.
Add a host, an endpoint, a token or a Prometheus target. No agent to roll out, no sidecar to inject — Orion polls from where you run it.
A monitoring daemon runs every check on a one-minute cron and writes the result. That is what makes the numbers move; nothing else writes check results.
The overview gives a single verdict, the wall gives every check, and the boards give whatever your team keeps open during an incident.
Sign in and point Orion at the first thing you want watched. It starts checking on the next minute.