I build the pipeline, ship the site,
and fix the data.
CI/CD that deploys itself. Websites that actually reach the internet. Complete ETL pipelines that finish. Datadog, Grafana and ELK that tell you the truth. Tell me the problem, the budget and the deadline — you get a scope and a price within one business day, or a straight no.
No sales sequence. One reply, from me, with a real number in it.
Monitor alert· Datadog · prod
Consumer lag is high on orders-consumer
kafka.consumer_lag > 200k for 15 min · prod
- Build
- Test
- Deploy
- Smoke
Three real shapes of work — a Kafka consumer that fell behind, a deploy that stopped shipping, a Datadog bill that tripled — from the alert to the fix landing in production.
Things I have actually shipped.
Not a logo wall. Each of these is live — click through and see it working.
This site’s own delivery pipeline
Push to main → tests → one image built and pushed to GHCR → the VPS pulls it → smoke test → automatic rollback if it fails. A database dump gates every deploy. Nobody SSHes in; a deploy is a merge and about ten minutes.
- GitHub Actions
- Docker
- PostgreSQL
A document conversion engine that never trusts the file
Markdown, DOCX, PDF, HTML and more through one reader × writer matrix. Anything uploaded is rendered inside an isolated container on its own network, reached only through a spool volume — never HTTP — and CI fails if that isolation ever loosens.
- Python
- Docker
- FastAPI
A SQL, Spark and Python assessment platform
Every problem runs in the visitor’s browser — SQLite compiled to WebAssembly for SQL and Spark, Pyodide for Python — with the grader owning every score. No execution backend to secure, scale or pay for.
- Python
- Apache Spark
- React
Datadog cost modelling and a monitoring library
A calculator that turns hosts, custom metrics, logs and APM spans into the bill Datadog will actually send, plus a course and a blog hub on the same subject — the same work I do inside a client’s account, published.
- Datadog
- Grafana
- Prometheus
Client work is under NDA more often than not, which is why the evidence here is my own. Reviews from clients appear below as they are approved — how that works.
This is what your next deploy looks like.
One push to main. Tests run, an image is built once and pushed, production pulls it, a smoke test confirms it is alive, and a rollback stays armed the whole time. Nobody SSHes into anything. Nobody deploys on a Friday holding their breath.
- Checkout & setuprunning…
- Lint & testqueued
- Build imagequeued
- Push to registryqueued
- Deployqueued
- Smoke test & rollback guardqueued
$ git checkout main
Before
A deploy is a person, a runbook and forty minutes of hoping.
After
A deploy is a merge. Three and a half minutes, unattended.
If it breaks
The smoke test fails, the old image comes back, and you get a message.
Six things I do properly.
Not a menu of everything — a list of what has actually been run in production, more than once, at the far end of a pager.
Not a chatbot bolted onto a dashboard — a governed AI analyst your business users can actually ask questions and get real answers from.
- Genie Agent + Unity Catalog, wired to the tables people are allowed to see
- An agent embedded in your own SaaS or internal app, not a side tool
- Multi-agent platforms on MCP, Genie and Lakebase for more than one domain
- Legacy BI (Power BI, Tableau) migrated onto an ask-anything analytics layer
- Governance underneath it: access control, evaluation, monitoring, audit trail
- Databricks
- Python
- FastAPI
Every push builds, tests and deploys itself — and rolls back on its own when it should not have.
- GitHub Actions pipelines, from scratch or rescued
- Docker images built once and promoted, not rebuilt per environment
- Zero-downtime deploys with an armed rollback
- Secrets out of the repo and into the platform
- Kubernetes, ECS or a single VPS — whichever is honestly right
- GitHub Actions
- Docker
- Kubernetes
- Terraform
Built new, or taken over from whoever left. Shipped to a real domain with TLS, caching and monitoring already on.
- Next.js, React and FastAPI builds
- Deploying an existing site that has never made it live
- Migrations off a host that no longer works
- Performance, Core Web Vitals and SEO fixes that hold
- Payments, auth and email that actually deliver
- Next.js
- React
- FastAPI
- PostgreSQL
A complete pipeline — source, transform, table, dashboard — built and running, not a proof of concept.
- End-to-end ETL and ELT, from extract through to a dashboard someone opens
- Airflow DAGs with retries, backfills and no silent skips
- dbt models, tests and lineage on top of the warehouse
- Spark on EMR or Databricks — runtime and cost brought back down
- Unity Catalog, Genie and Iceberg tables that stay queryable
- Apache Airflow
- Apache Spark
- dbt
- Databricks
Streams that stay caught up, and stay caught up after the next replay.
- Kafka consumer lag diagnosed and drained
- Kinesis shard, checkpoint and throttling problems
- Structured Streaming and CDC into the lakehouse
- Exactly-once semantics that survive a real replay
- Schema drift caught before it reaches the table
- Apache Kafka
- Apache Spark
- AWS
Dashboards that answer the question during an incident, and alerts that only wake someone who can act.
- Datadog dashboards, monitors, APM traces and SLOs, per service
- Grafana and Prometheus — self-hosted or Grafana Cloud
- ELK log pipelines, index lifecycle policies and shard sizing
- Alert noise cut back to the pages that need a human
- A monitoring bill traced to the tag that is actually causing it
- Datadog
- Grafana
- Elasticsearch
- Prometheus
If a person does it every week from a runbook, it should be doing itself by next week.
- Release and deployment automation
- Scheduled reports and data exports
- Infrastructure as code — Terraform, Compose
- Alerting that pages on causes, not symptoms
- Glue scripts and internal tools nobody else will own
- Terraform
- GitHub Actions
- Docker
I don't sell “a Genie Agent.” I sell the AI analyst it becomes.
“I'll build an enterprise AI data analyst that lets your business users ask questions about your company's data and get back governed answers, SQL, visualizations and reports” — that is the product. Genie, Unity Catalog and agent frameworks are how it gets built, not what it's called on the invoice.
Ask-your-data AI analyst
$$$$$Business users ask a question in plain English and get back governed SQL, a result set and a chart — scoped to the tables they are actually allowed to see.
Databricks · Unity Catalog · Genie · SQL
Enterprise data AI agent
$$$$$An agent that reasons over your metrics, flags anomalies unprompted, and can act — open a ticket, page a team — not just answer when asked.
Genie · Agent framework · Python
AI analyst embedded in your product
$$$$$The same ask-your-data experience, shipped as a feature inside your own SaaS or internal tool instead of a separate app your users have to context-switch to.
Genie API · FastAPI · React
Multi-agent analytics platform
$$$$$Several specialised agents — one per domain — sharing context and handing off to each other, backed by a live operational data store rather than a nightly export.
Agent orchestration · MCP · Genie · Lakebase
A full-time AI data analyst for the business
$$$$$One system that answers finance, ops and product’s questions off a shared semantic layer — so three teams stop getting three different numbers for the same metric.
Databricks · Genie · Semantic layer
Legacy BI → AI analytics migration
$$$$$Retire the dashboards nobody trusts and move onto a governed, ask-anything analytics layer, without breaking the reports finance already depends on.
Databricks · Power BI / Tableau · Genie
Production agent platform, governed
$$$$$The layer underneath the agents: access control, evaluation and monitoring, and an audit trail — the part that makes it safe to point an agent at production data.
Unity Catalog · Security · Evaluation · Monitoring
Not sure which of these fits, or want to talk through the architecture first? Book a call and describe what your business needs to ask its own data.
I build the whole ETL, not a proof of concept.
Sources through transform to a dashboard someone will actually open — stood up in a day, then hardened: retries, backfills, idempotent reruns and schema drift caught before it lands. Airflow and dbt where the orchestration belongs, Spark where the work is, Unity Catalog and Genie for the ones already on Databricks, plain Spark and Iceberg for the ones who are not. Kafka lagging or Kinesis throttling on top of it is a Tuesday.
One pipeline, end to end
- SourcesKafka · Kinesis · S3 · Postgres
- IngestStructured Streaming · CDC
- TransformSpark on EMR / Databricks
- ServeUnity Catalog · Iceberg
- DashboardGenie · BI · alerts
Consumer lag
caught up- Peak lag
- 271k
- Drained in
- 41 min
Rebalancing storm, a poison-pill partition, a consumer that quietly stopped committing — the cause is usually not where the alert points. We find it, then fix the thing that let it happen.
Datadog, Grafana and ELK — set up, cleaned up, or paid down.
Most teams already have one of these. The problem is rarely that it is missing — it is that nobody trusts the dashboards, the alerts fire at 3am for something nobody can fix, and the invoice has quietly tripled. I take over whichever one you run, or stand up the right one from scratch.
Datadog
Monitors that page a human only when a human is needed — and a bill that stops climbing.
- Dashboards and monitors built per service, not per metric
- Custom-metric cardinality traced to the tag causing it
- APM and distributed tracing wired through the stack
- SLOs, burn-rate alerts and a real on-call rotation
- Log pipelines with indexes scoped to what gets queried
Grafana & Prometheus
Self-hosted or Grafana Cloud, with exporters and recording rules that survive a scrape restart.
- Prometheus exporters, scrape configs and recording rules
- Dashboards someone opens during an incident, not once a quarter
- Alertmanager routing, grouping and inhibition that hold
- Loki for logs and Tempo for traces alongside the metrics
- Retention and cardinality tuned before storage becomes the problem
Elastic stack (ELK)
Logs that are searchable in seconds, on a cluster that is not quietly eating your budget.
- Logstash and Beats pipelines with parsing that does not drop fields
- Index templates, ILM policies and sane shard sizing
- Kibana dashboards and saved searches per team
- Cluster health, hot-warm tiers and rebalancing fixed
- Migrations to OpenSearch, or off it, without losing history
Pages per week, before and after a noise pass
Before
After
An alert nobody acts on is not monitoring — it is training your team to ignore the one that matters. Thresholds get tied to symptoms a user would notice, and everything else becomes a dashboard.
The bill is part of the job
Observability bills rarely grow because you added hosts. They grow because one tag went high-cardinality and nobody noticed for a quarter.
Custom metrics are billed per unique tag combination, so a single user_id or request_id tag can multiply one metric into hundreds of thousands. Every engagement starts by finding those before it touches a dashboard.
The delay is usually the process, not the problem.
Most of the time an issue is not hard — it has just been sitting in a queue behind a discovery call, a statement of work and someone's sprint boundary. I took those out.
1 day
Application broken?
Most defects are diagnosed and fixed inside a single working day. If it is bigger than that, I tell you on day one — not on day five.
1 day
A complete ETL
A working end-to-end pipeline — source to table to dashboard — stood up in a day, then hardened. No three-week discovery phase.
< 24h
Scope and price
Every enquiry gets a real answer within a business day: what it takes, what it costs, or a straight "this is not a fit".
From your enquiry to work starting.
You tell me the shape of it
Four short steps: what you need, how urgent it is, what the budget band is, and what is actually going wrong. It takes about two minutes.
I read it properly and reply
I read every word myself. You get back a scope, a price and a first opinion on the cause — usually within one business day.
A call, only if it earns its place
Thirty minutes, on the record, with the technical questions already written down. No discovery call whose only purpose is another call.
I do the work
You approve the price, then I build it and hand it over — with the cause explained, so the same thing does not come back next quarter.
- I turn work down. If it is outside what I have run in production, you get told on day one instead of paying me to learn.
- The price is fixed before any paid work starts, and it does not move because the work turned out harder than I guessed.
- One reply, from my own address. No sequence, no newsletter, and your details are never passed on.
Tell me what you need.
Four steps, about two minutes. The first two are taps. I ask about budget and timing up front for one reason: so the reply you get is a real answer instead of a request for another meeting.
Want your own team to be able to build these instead of hiring it out? There's a free course — Forward Deployed Engineer.

Not ready to scope a project?
Book a free 1:1 call to talk it through first.
No budget or timeline questions — just describe what's going on and I'll tell you straight whether it's a fit.
Questions people ask first
How fast can you actually start?
For an emergency — production down, a deadline today — usually within hours of your enquiry landing, because those are graded to the top of the inbox. For scoped work, a few days, depending on what is already in flight. You will be given a real date, not a "soon".
Why do you ask for a budget band before we have even spoken?
Because it decides the shape of the answer rather than the price of the hour. A small budget gets an honest "here is what fits inside it", which is a useful answer. A budget that cannot cover the work gets told so immediately, which is a better answer than a call that ends the same way an hour later. Nobody is asked for a figure they have not approved — it is a band, and "just exploring" is one of the options.
Can you take over a project someone else started?
Yes, and it is a large share of the work. Half-finished repositories, a site that was built but never deployed, a pipeline whose author has left — all normal. The first thing you get is an honest read on what is salvageable and what is cheaper to rebuild.
Do you need access to our production systems?
Not to begin with. Most problems can be diagnosed from logs, an error message and a description of what changed. When access is genuinely needed, we agree the narrowest scope that works, through credentials you issue and can revoke the moment the work is done.
Can you really build a complete ETL in a day?
A working one, yes — extract, transform, load, and a dashboard at the end of it, running on your data rather than a sample. What a day does not buy is a hardened one: retries, backfills, idempotent reruns, schema-drift checks, alerting and cost controls come after, once you have seen it work and confirmed the shape is right. Building it in that order means you find out on day one whether the source data actually supports what you wanted — which is where most three-week discovery phases end up anyway.
We already pay for Datadog. Can you work with what we have?
That is the usual case, and it is the cheaper one. Whichever of Datadog, Grafana, Prometheus or the Elastic stack you already run, I work inside it rather than proposing a migration — a rip-and-replace mostly moves the problem to a new vendor and costs you your history. Two things get looked at first: which alerts fired in the last month without anyone acting on them, and which tags are multiplying your custom-metric count. Those two answers often cover the cost of the engagement on their own.
Is this a one-off fix or ongoing support?
Either. Some clients want one specific thing fixed and then to be left alone; others keep a few days a month on retainer so the next thing has somewhere to go. You pick the shape in the first step of the form, and it changes what I quote.
Do you work with clients outside India?
Yes — most of it is. Everything is remote (email, shared repositories, screen share) and invoicing works in rupees or dollars. Time zones have not been a problem; overlapping hours are agreed before the work starts.
What happens if you cannot solve it?
You are told that, and you pay nothing for the attempt. Where I can, I point you at who or what would help instead. Turning work down costs me one enquiry; taking work I cannot finish costs a reputation.
Just one thing broken and nothing bigger behind it? Send it to /fix instead — same person, shorter form, fixed price.
