Observability platform by Tellestia

Know before
they tell you.

Orion polls every service you run — pings, API endpoints, tokens, logs, metrics and clusters — and puts the state of all of it on one screen. When something breaks, the board shows it before the first support ticket lands.

  • Self-hosted
  • No agent to deploy
  • Light and dark
  • Keyboard-first
Live

Fleet verdict

1 service down

Uptime

99.97%

Mean RTT

50.3ms

Checks run

1,284,960

Fleet latency · last 20 min

peak 63.0 ms

Live stream

  • pingedge-sin responded in 41.2 ms
  • apiGET /v2/catalogue/search · 200 · 96 ms
  • apifraud-score p95 above 400 ms threshold
  • pingkafka-03 recovered after 2 failed checks
  • apipartner-gateway returned 503 · 3rd consecutive
  • loki18,402 lines ingested from prod/checkout

Signal wall · 32 checks

edge-fraedge-sinedge-iadcore-rtrpg-pay-01pg-pay-02kafka-01kafka-02kafka-03redisobjectsldapsmtpvpn-hqcdn-shieldk8s-apilokicheckoutauthoriseoauth2catalogueinventoryshippingprofilepricingfraudnotifyhooksreportspartnersflagsbilling
Illustrative interface · sample estate

Reads the systems you already run

PrometheusLokiKubernetesNGINXMongoDBPostgreSQLKafkaRedisWSO2JiraGrafanaDockerOpenSSHElasticsearchRabbitMQConsulPrometheusLokiKubernetesNGINXMongoDBPostgreSQLKafkaRedisWSO2JiraGrafanaDockerOpenSSHElasticsearchRabbitMQConsul
Platform

Four layers of your estate, one platform

Most teams watch availability in one tool, logs in a second, metrics in a third and clusters in a fourth — then reconcile them by hand at two in the morning. Orion is the one screen those four were supposed to be.

Availability

Pings, endpoints, tokens

Reachability, status codes, response times and credential expiry — every target checked on a fixed cadence, every minute, from the moment you add it.

  • ICMP and TCP targets
  • HTTP checks with assertions
  • Token expiry countdowns
Telemetry

Metrics and logs, one screen

Prometheus for metrics, Loki for logs, queried side by side instead of in three consoles with three logins and three mental models.

  • PromQL and LogQL surfaces
  • Live tail with level filters
  • Retention you control
Infrastructure

Clusters, nodes, gateways

Kubernetes workloads, node vitals, NGINX gateways and WSO2 traffic, each with its own view and all rolling up into the same fleet verdict.

  • Pod, node and namespace views
  • Gateway error-rate breakdown
  • Asset discovery sweeps
Intelligence

Answers, not just charts

Point Orion at a failing window and it writes the root cause from the log lines themselves, finds the incident it resembles, and turns plain English into a query.

  • Root-cause drafting
  • Incident similarity search
  • English to LogQL

60 s

check cadence

every monitor, every minute

5

signal types

ping · API · token · log · metric

1

screen

the whole estate, one verdict

100%

self-hosted

your data never leaves your network

Capabilities

Built for the ten minutes that matter

Dashboards are easy to make pretty and hard to make useful. Every surface in Orion is designed for the moment something is wrong and someone is trying to find out what.

Availability

Know the moment something stops answering

Every check runs on its own schedule and writes its result. Orion keeps the pass and fail counts, so a monitor that is technically up but failing one check in twenty is impossible to hide.

  • One cell per monitored thing, so the whole estate reads in a glance
  • Reliability shown as a proportion of real checks, never a decorative bar
  • Degraded is its own state — not rounded up to healthy or down to failed
See the uptime board

Endpoint reliability · 24 h

60 s cadence
  • checkout148 ms
  • authorise212 ms
  • fraud-score243 ms
  • catalogue96 ms
  • partner-gw—
  • flags19 ms
Telemetry

The log line and the metric spike, together

When latency moves, the reason is usually in the logs from the same two minutes. Orion queries Loki and Prometheus from one screen so you stop copying timestamps between tabs.

  • LogQL and PromQL, with the query kept in the URL so it is shareable
  • Level, service and environment filters that survive a page reload
  • An empty source says which service to connect, never a red error
Open the log explorer
{app="checkout", env="prod"} |= "503"
  • 14:02:10INFOorder accepted id=8841 amount=129.00
  • 14:03:11INFOauthorising card ending 4242
  • 14:04:12WARNretry 1/3 upstream=partner-gateway
  • 14:05:13ERRORupstream 503 after 3 attempts
  • 14:06:14INFOorder queued for manual review
18,402 lines matched · 340 ms
Orion Intelligence

It reads the logs so the on-call engineer doesn't have to.

Orion drafts a written root cause from the log lines around a failure, cites the ones it used, and tells you which past incident this one resembles. Ask it a question in English and it writes the query.

  • Root causeDrafted from the log window, with citations.
  • Similar incidentsMatched against everything you have resolved.
  • English to LogQLDescribe it; Orion writes the query.
  • Anomaly hintsLog patterns that are not normal for this service.

Runs against a self-hosted model — your logs stay inside your network.

Root cause · partner-gatewaydrafted 14:06

What happened. Every request to partners.orion.internal/v1/orders has returned 503 since 14:02. Three consecutive checks failed.

Why. The upstream began refusing connections two minutes after checkout-7d9f was OOMKilled and restarted — the connection pool was not re-established.

12 log lines citedINC-2291 looks similarconfidence: high
How it works

Watching something within the hour

Orion runs where your services run. There is no agent to roll out and nothing to install on the things being watched.

  1. 01

    Point it at something

    Add a host, an endpoint, a token or a Prometheus target. No agent to roll out, no sidecar to inject — Orion polls from where you run it.

  2. 02

    The daemon takes over

    A monitoring daemon runs every check on a one-minute cron and writes the result. That is what makes the numbers move; nothing else writes check results.

  3. 03

    Watch one screen

    The overview gives a single verdict, the wall gives every check, and the boards give whatever your team keeps open during an incident.

OverviewOne verdict for the fleet
DashboardsBoards you assemble yourself
AlertsOnly for what actually failed
TokensExpiry before it bites

Put the whole estate on one screen.

Sign in and point Orion at the first thing you want watched. It starts checking on the next minute.