Written by Huifer, solo developer and maintainer of TanStack Ship. I run performance monitoring and product analytics side by side across 12+ TanStack Ship SaaS apps on Cloudflare Workers — apps serving tens of thousands of MAU that I built and I page myself when something breaks. The split: edge metrics, structured logs, and tail-worker traces on the performance side; self-hosted PostHog for analytics; all joined through a shared request id propagated from browser to Worker to database. Every threshold, query, and alert here is one I ship. The bias toward TanStack Ship is named up front — it is my paid product.
Verified sources: Cloudflare Workers Analytics Engine · Cloudflare Workers Observability · Web Vitals · OpenTelemetry · PostHog open-source product analytics · Google SRE Workbook — SLO chapter
Last updated: 2026-07-12 · Changelog
TL;DR: Performance monitoring and product analytics are not the same discipline, but they fail the same way when run in isolation — performance tells you a metric moved, not what users did; analytics tells you what users did, not whether the system carried them. The unified stack I run on Cloudflare Workers in 2026 has five performance layers plus a product analytics layer joined by request id, wrapped in SLIs and SLOs that page only on user-visible degradation. This guide covers the five-layer telemetry stack, the product analytics layer, SLI/SLO math, cost economics at three app sizes, dashboards, and the weekly review ritual that keeps the system honest.
Why Monitoring and Analytics Belong in the Same Conversation
Most SaaS teams run monitoring and analytics as two separate stacks — Datadog for one, Mixpanel for the other — drawing two dashboards and never joining the data. The result is a fragmented picture where you can answer half of every question.
The blind spots of monitoring alone
A performance stack answers is the system healthy? with metrics, logs, and traces. It detects that latency regressed on /api/checkout, that the 5xx rate spiked after a deploy, or that a database query now scans 40,000 rows instead of 40. It is silent on whether the spike hit one user or ten thousand, whether it sat on the conversion path or an unused admin screen, and whether a new feature triggered the regression.
I learned this the expensive way. In Q1 2026, a TanStack Ship billing module deployment pushed p95 checkout latency from 320 ms to 780 ms. The alert fired, the rollback ran, I felt good — until the analytics dashboard showed conversion had been down 11% for the four hours the regression was live. Monitoring had told me something was wrong; it could not have told me what was at risk.
The blind spots of analytics alone
Product analytics answers what did users do? — signups, activations, funnel drop, retention, feature adoption. It tells you trial-to-paid conversion dropped four points last week. It is silent on why — bug in the upgrade flow, a slowdown that crossed a usability threshold, or an upstream OAuth outage.
A SaaS team I advised ran two months on analytics alone. Their conversion drop turned out to be a slow OAuth handshake that added four seconds to first login. They had no traces for the OAuth path and found it by reading logs of a single support ticket.
What "unified" actually means
Unified does not mean one vendor. It means a shared identifier — a request id the browser sends, the Worker logs, the analytics event carries, and the database query attaches. Once the same id shows up in both dashboards, you can answer the question neither side can answer alone.
In practice that is a UUID propagated from client to every server event, written into the Analytics Engine row, attached to the PostHog event, and stamped onto the D1 query. Joining on that id turns two stacks into one.
The Five-Layer Performance Stack I Run
Performance telemetry divides into five layers. Each has a job; none substitutes for another.
Layer 1: Edge metrics on Workers Analytics Engine
Aggregated metrics live on Workers Analytics Engine — a time-series database for high-cardinality edge workloads: write-heavy, SQL-queryable, cheap at the volumes I operate.
-- p95 latency per route over the last 6 hours
SELECT
blob1 AS route,
int2 AS status_class,
quantileCont(0.95)(double4) AS p95_ms
FROM response_metrics
WHERE timestamp > NOW() - INTERVAL '6' HOUR
GROUP BY route, status_class
ORDER BY p95_ms DESC
LIMIT 25;
Every Worker writes one row per response with route, status class, duration, and dimension columns (deploy id, region, request id). One dashboard per app pulls this query every five minutes. The same data feeds the SLI computation that drives alerting.
Layer 2: Structured logs in Workers Logs and Tail Workers
Metrics tell you a number moved; logs tell you which requests contained it. Workers Logs captures console output and exception traces with request correlation built in. I log structured JSON, so I can grep with jq and pipe into incident reviews.
A Tail Worker holds the per-request detail — request log with PII redacted, headers minus cookies, body shape but not body values, and the request id that ties everything together. I sample at 1-in-10 by default; I bump to 100% for any deploy I am debugging.
Layer 3: Distributed traces via Tail Workers and OpenTelemetry
Traces answer which database query, in which downstream call, on which deploy, caused this regression. The OpenTelemetry pattern works on Workers with one caveat — there is no in-process agent, so I export OTLP from a Tail Worker.
// tail-worker.ts — emits an OTLP span per sampled request
export default async function (events: TraceItem[], env: Env) {
const spans = events.map((e) => ({
traceId: e.requestId,
name: `${e.event.request.method} ${e.event.request.url}`,
startTimeUnixNano: e.event.cp,
endTimeUnixNano: e.event.cp + (e.event.wallTimeMs * 1_000_000),
attributes: {
'http.method': e.event.request.method,
'http.route': e.event.scriptName,
'http.status_code': e.event.response?.status,
'deploy.id': env.DEPLOY_ID,
},
}));
await fetch(env.OTLP_ENDPOINT, {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ resourceSpans: [{ spans }] }),
});
}
I have not benchmarked this above 5k spans/sec sustained — at the volumes I operate it has never been the bottleneck. If you cross that line, switch to a sampling tail worker.
Layer 4: Real-user monitoring with web-vitals + beacon
Server-side metrics tell you what your edge saw. RUM tells you what the user's device actually experienced. The two diverge — your edge can return a 200 in 80 ms while the user sees a six-second LCP because a third-party script blocked the main thread. My RUM guide covers the implementation; the short version is web-vitals in the browser, a sendBeacon to a Worker endpoint, and the same request id flowing into Analytics Engine tagged source=rum.
Sample rate is the trap. At 100% you burn CPU and quota on telemetry; at 1% you lose visibility into low-traffic tenants. I sample at 10% globally, 100% on error responses, and 100% during any active incident.
Layer 5: Synthetic checks on cron-triggered Workers
RUM only fires when real users hit your routes. Synthetic checks fire whether anyone is looking. A cron-triggered Worker hits the most critical paths every minute — login, signup, the upgrade flow, the pricing page — from three regions. Latency regressions show up here first.
I learned to value synthetic checks during a 2025 D1 write-contention incident: the synthetic checker paged me 90 seconds before the first customer-facing impact.
The Product Analytics Layer That Makes Performance Useful
Performance tells you the system moved. Product analytics tells you which user took which action. Together they answer the question that pays the bills.
Event collection without breaking the bank
I run PostHog self-hosted on a small VM. It captures the events that matter — signups, activations, upgrades, churn, feature usage — and joins to performance data on request id. Cost is a function of monthly tracked users, not events — the right shape for a SaaS where each user generates dozens of events.
A common mistake: instrumenting everything. I keep the client event surface to about 12 named events. The rule: if an event does not change a decision I would otherwise make, I do not collect it.
Funnel, retention, and feature usage
Three analytics views earn their keep on every SaaS dashboard:
- Activation funnel — the four-to-six steps from landing page to the first moment the user perceives product value. A 5% drop is a regression.
- Cohort retention — the percentage of users from each signup week still active in week N. The shape tells you more about your product than any single metric.
- Feature usage heatmap — what percentage of paid users touch each feature. Below 5% is a removal candidate; above 40%, deeper investment.
Joining product events to performance traces
This is where the unified stack earns its name. Every PostHog event carries a request_id property. Every Analytics Engine row carries the same id. To find the slow checkout path that cost a paying customer,
// find the performance trace for users who cancelled
const slowCheckout = await analytics.query(`
SELECT * FROM response_metrics
WHERE timestamp > NOW() - INTERVAL '7' DAY
AND blob1 = '/api/checkout'
AND double4 > 500
AND blob3 IN (
SELECT distinct_id FROM posthog_events
WHERE event = 'subscription_cancelled'
AND timestamp > NOW() - INTERVAL '7' DAY
)
LIMIT 100;
`);
That query is the difference between "our checkout got slow last week" and "checkout got slow for the cohort that cancelled, and here are the request ids to inspect." It is the question neither side can answer alone.
SLIs, SLOs, and Error Budgets That Actually Page the Right People
Without SLIs and SLOs, alerts fire on every metric that wiggles. With them, alerts fire only when users are feeling pain you promised not to inflict.
Choosing SLIs users feel
An SLI is a ratio: good events divided by valid events. The four I track:
- Availability — 2xx responses divided by all responses, excluding redirects.
- Latency — responses under 300 ms divided by all responses, at p95.
- Correctness — successful writes divided by all write attempts.
- Freshness — successful background-job runs divided by scheduled runs.
The discipline is what the user feels. Queue depth is a signal; it is not an SLI unless the user can perceive it.
SLO targets that survive contact with reality
A target you cannot hit is theatre. A target too easy to hit is not a contract. The targets I run:
- Availability: 99.9% over a rolling 28-day window (~40 minutes of budget per month)
- Latency: 95% of requests under 300 ms at p95
- Correctness: 99.99% on write operations
- Freshness: 99% on background jobs
I review the targets quarterly. Apps burning less than 50% of budget are candidates for tightening; apps burning more than 80% are candidates for real reliability investment.
Alert routing and burn-rate alerts
The point of an SLO is the error budget — the badness you have allowed yourself over a window. Burn-rate alerts fire when you are consuming budget faster than sustainable. I run two:
- Fast burn — 2% of budget in one hour. Pages me directly.
- Slow burn — 5% of budget in 24 hours. Slack notification, no page.
Burn-rate alerting replaced threshold alerting on every app I run. The shift took about a month to tune and produced a roughly 4× reduction in pages per quarter.
Cost Economics: Three SaaS Sizes, Two Stacks
The most-asked question I get is what does this stack actually cost? It depends on volume, but the shape is predictable.
Tiny app — under 10k MAU
The Cloudflare-native stack is essentially free. Workers Analytics Engine includes 100,000 events per day on the paid Workers plan; PostHog self-hosted on a $5/mo VM handles analytics. The third-party stack — Datadog plus Mixpanel — costs roughly $80-200/mo at the same volume. The free stack has obvious appeal.
Mid app — 10k to 100k MAU
The free stack starts to show friction here. Analytics Engine writes scale linearly; at 100k MAU you are likely hitting limits on the included quota and need either the paid add-on or aggressive sampling. Total monthly run cost: $30-200.
The third-party stack at this size is $400-1,500/mo. The economic case for the Cloudflare-native stack is strong — but the operational case is a real tradeoff. I have one app in this band on the third-party stack because the on-call burden of self-managing traces was not worth the savings.
Larger app — over 100k MAU
At this size the two stacks converge in cost, and the third-party stack usually wins on features. Datadog's APM, Honeycomb's BubbleUp, New Relic's error tracking are mature products with years of investment. Total monthly run cost: $200-1,500 (Cloudflare-native) versus $1,500-6,000 (managed stack).
When the managed observability stack wins
Three rules of thumb:
- If your engineering team is more than two people, the collaboration features of Datadog or Honeycomb often pay for themselves.
- If your trace volume exceeds roughly 5,000 spans/sec sustained, the managed stack's indexing pipeline is more reliable than anything I would build myself.
- If you operate across multiple cloud providers, the third-party stack is the only way to get a unified view.
For solo-dev SaaS apps on Cloudflare Workers — which is what TanStack Ship is designed for — the Cloudflare-native stack is the better default.
Dashboards, Alert Routing, and the Weekly Review
Telemetry you do not look at is wasted telemetry.
Three dashboards, not thirty
Every TanStack Ship app I run has three dashboards:
- Production health — p95 latency, error rate, request volume, deployment marker overlay. One screen, glanceable.
- User funnel — signup → activation → upgrade → retention. Weekly view.
- SLO status — current 28-day error budget remaining for each SLI, color-coded green/amber/red.
Everything else is a drill-down, not a top-level dashboard. If a regression does not show up on one of the three, it is not worth a page.
Alert severity ladder
Four levels:
| Severity | Routing | Examples |
|---|---|---|
| P0 | Phone page | Fast burn SLO alert, complete outage |
| P1 | Slack + phone | Sustained error rate above threshold, deploy regression |
| P2 | Slack | Slow burn SLO alert, single-route latency regression |
| P3 | Email/digest | Capacity warnings, expired certificates |
The mistake I see most often is too many P1s. If everything is P1, nothing is. The ladder forces the question what would I actually do at 3am? — and if the answer is nothing until morning, it is not a P1.
The Monday morning review ritual
Every Monday at 9am I spend 30 minutes reviewing:
- SLO burn rate for the previous week
- Top three latency regressions by user impact
- Top three conversion funnel changes
- Any P0/P1 incidents and their postmortem status
That is the entire ritual. I have not missed a customer-visible regression in two years — the weekly review catches what the alerts miss.
The Stack as a Single System
Monitoring and analytics fail the same way when run in isolation: each answers half of every question. Joined by a request id and operated with SLIs/SLOs as the contract, they become a single system — one that pages only on user-visible degradation, joins telemetry to business outcomes, and costs a solo dev time but not a managed-platform budget.
The stack I run in 2026 is not exotic. It is the boring consequence of having been paged at 3am too many times: five performance layers, one analytics layer, one shared identifier, SLOs, three dashboards, and a Monday morning review. If you ship a SaaS on Cloudflare Workers and you are not running something close to this, the TanStack Ship boilerplate ships with the telemetry layer pre-wired — so you start where I had to earn my way to.
Next steps: read the observability deep-dive for the three-pillar implementation detail, the RUM guide for the client-side instrumentation, and the analytics stack setup for the PostHog configuration I run. Compare TanStack Ship's monitoring defaults against Shipfast and Makerkit before you decide which boilerplate carries your telemetry for you.