Now in private access · AWS · GCP · Azure · Kubernetes

The on-call rotation, minus the pager.

Always-on cloud ops agents that detect configuration drift, self-heal incidents around the clock, and ship a curated weekly reliability digest to your engineering team — without locking you to a single cloud or charging enterprise AIOps prices.

Built by Fred Bitenyo — Senior DevOps Engineer, former pager.

mendhelm · /ops/live
LIVE
[ok] drift · eu-west-1 · iam/policy:mendhelm-prod04:11
[ok] deploy · api · 3.42.1 → 3.42.204:09
[warn] heal · k8s/checkout · pod 7 loop · restarted03:58
[ok] patch · base image · sha pinned03:41
[ok] drift · gcp · service-account:mendhelm-stg03:22
[ok] audit · 247 actions · 0 escalations03:00
agents: 4 online · esc: 0 · drift: 2 cleared / 0 pending
Operates acrossAWSGCPAzureKubernetes

Capabilities

Four jobs. Run continuously. Signed off on Monday.

Mendhelm is one product, four operators. Each one runs the same way your platform team would — diff first, dry-run, promote, audit — except it doesn't sleep, doesn't take holidays, and never goes a week without reporting back.

OperatorWhat it doesState
CND-01Drift detection

Continuous diff against declared and provisioned state. Catches config drift, IAM policy drift, label drift, and tag drift before they cascade into incidents.

OK
CND-02Self-healing remediation

Agents open, triage, and close incidents autonomously — restart the stuck deployment, roll back the bad release, drain the unhealthy node, raise the limit, re-attach the volume.

WARN
CND-03Patching and config rollout

Push pinned versions, rotate secrets, apply Terraform deltas, patch base images. Each rollout is sandboxed, then promoted — never the other way around.

OK
CND-04Weekly reliability digest

A curated engineering-team report every Monday: what drifted, what healed, what an on-call human should still sign off on. Not a dashboard. A summary a director can act on.

OK

How it works

Connect once. Operate the rest of the year.

Implementation is a single OIDC connector and a one-page scope config. There is no agent to install, no sidecar to deploy, no "observability tax" to pay.

Read the agent loop deep dive →
step / 01✓ ready

Connect once, scoped tight

Read-only by default. Agents assume least-privilege scopes per environment, per service, per region. You grant exactly the blast radius you trust us with.

step / 02✓ ready

Watch and learn

Mendhelm baselines your environments over the first 72 hours — normal traffic shapes, deployment cadences, idle baselines. After baseline, drift becomes signal instead of noise.

step / 03✓ ready

Operate, audit, report

Agents remediate with dry-runs first, then promote. Every action is logged with prompt, scope, diff, and outcome. Weekly digests roll the work up for the team.

Security & controls

What runs behind every agent action.

Giving an agent permission to change production is a trust motion. We earn it by making the controls explicit, the audit trail complete, and the off-switch always one click away.

Request the security brief
  • Layered permissions and per-environment scopes
  • Least-privilege IAM roles, rotated weekly
  • Dry-run mode that defaults to on for any new action class
  • Per-agent kill switch, per-environment kill switch
  • Full audit trail: prompt, scope, intended diff, actual diff, rollback
  • Change windows and freeze periods honored verbatim

audit.last_24h · 247 actions · 0 escalations · 0 anomalous scopes

Weekly reliability digest

A report your director will actually read.

Every Monday, Mendhelm sends a curated engineering-team report — not a raw dashboard. What drifted. What healed. What an on-call human still owns. Reliability framed as a product engineering output your leadership and rotation can act on.

monday · 09:00 local · your inbox

Reliability digest · week 34

Mendhelm · weekly

ready

drifted · cleared

6 IAM policy drifts, all auto-patched in under 4 minutes

The weekly baseline turned up six drifted assume-role policies on themendhelm-prodaccount. Each was diffed against the last approved manifest, dry-run, then promoted. No engineer was paged.

healed · self

Checkout pod wanted to page on-call. We didn't let it.

A 7-pod crash-loop in checkout-prod cleared after a rolling restart and a softened readiness probe that won't mask a real outage next time.

escalated · human

Auth service ahead of cost forecast — needs a human this week

Auth spends running 14% ahead of the rolling forecast. Mendhelm refuses to scale down further (you asked us not to) — flagging for the platform owner.

next digest · mon 09:00 · delivered to eng-leads@yourcompanyRead the full sample →

Talk to Fred

Replace the on-call rotation with agents you'd actually trust.

Mendhelm is in private access. Founder Fred Bitenyo takes the first call personally — what you run, what hurts, whether agents are even right for you yet. No SDR.

Email mendhelm-3@polsia.appresponse within 24h · no automation · no SDR