RedCyferBusiness technology Call (870) 876-3016

Work / case study

In-house observability

A fully self-hosted monitoring platform built to replace a third-party SaaS. Every server, website, database, storage node and backup job on one dashboard, on infrastructure the company owns.

Replaced
Datadoga third-party SaaS
Environment
HybridAWS and on-premises
Hosted
In-houseinside the private network
Alerts
24/7to Slack, on trouble and on recovery

The brief

Like most companies, this one, a nationwide e-commerce operation running a hybrid AWS and on-premises environment, relied on a third-party SaaS (Datadog) to monitor its systems. It worked, but it meant a growing monthly bill and a continuous stream of operational telemetry flowing into someone else's cloud. The question was simple: what if we built our own?

The company isn't named here by agreement. It's the environment Chris runs day to day. A deeper walkthrough is available in conversation.

What was built

  • Metrics

    Live health of every server (CPU, memory, disk, network), plus load balancers, auto-scaling, CDN and cloud storage on one screen

  • Logging

    Fleet-wide logs in one searchable home, so an issue can be traced across systems in seconds

  • APM

    End-to-end request tracing that pinpoints slow pages and surfaces errors as they happen

  • Traffic

    Real-time web analytics with geographic and device breakdowns, plus bot detection

  • Alerting

    Configurable rules watching 24/7, posting to Slack on threshold crossings and on recovery

  • Security

    Continuous automated detection of suspicious activity, with early notification

  • Storage

    Direct health monitoring of an on-premises storage cluster

  • Backups

    Every backup run tracked and confirmed, with Slack and email summaries

Under the hood

Instead of a generic tool bent to fit the stack, the platform is the stack, watching itself. It was designed around how the environment actually runs.

What it runs on

  • Backend

    Node.js 22 and Fastify, with Python services for alerting and notifications

  • Data

    PostgreSQL / TimescaleDB for time-series metrics; ClickHouse for high-volume logs and traces

  • Frontend

    Dependency-light dashboard in vanilla JS and Apache ECharts. Fully self-hosted, no CDNs

  • Instrumentation

    OpenTelemetry for application tracing

  • Integrations

    AWS CloudWatch, Slack, the firewall and the storage cluster

  • Reliability

    Independent, isolated, self-restarting services. The monitor is not a single point of failure

  • Security-first

    Runs entirely inside the private network, VPN-gated, authenticated, with least-privilege read-only integrations

Why it matters

The data stays home. The bill stops growing with every host and log line. The tool fits the stack instead of the other way around. And when something breaks, the team finds out from their own platform, one screen they own outright. The thinking behind it is written up in own your telemetry.

Need software built around how you work?

Call (870) 876-3016, text (870) 641-5054, or send a note.