The stack I've built in production

Not side projects. Not tutorials. These run on real infra serving real users.

[LIVE IN PROD]

TraceLens — Autonomous Incident Investigation Agent

The thing that actually woke me up at 3AM before I built this. Now it wakes itself up, investigates, and files the Jira ticket. I just review it in the morning with coffee.

480×faster triage
40+tools
12+data sources
37tool schemas
Python 3.11FastAPISSEReact 18Azure OpenAI GPT-4oPostgreSQLAWS EKSNginx
[LIVE IN PROD]

Log Analysis Agent — S3 ReAct CLI + Browser UI

No vector store. No preprocessing. Just grep, awk, and zcat used by an LLM that actually understands what it's looking at. If that sounds stupid simple, that's because it is — and it works.

200k–30Mlines/service/day
Zeroingestion overhead
Bashas the toolset
PythonAzure AI FoundryRich Terminal UIprompt_toolkitReAct Loop
[PRODUCTION]

AlertFlow — Real-Time Alert Correlation

47 simultaneous alerts about the same thing isn't 47 problems. It's one problem and a terrible Tuesday. AlertFlow turns the storm into a queue.

185+services
Storm → Queuestructured triage
[REPLACED PAID SAAS]

Exception Clustering — Internal Sentry Alternative

Built an org-centric Sentry alternative from scratch. Then we cancelled our Sentry subscription. That felt good.

215+services
Groups · Dedupes · Ranksautomatically