Bellwether

Your customers already told you what to fix. Bellwether ships it, then checks it worked.

Bellwether reads your tickets, errors and analytics, finds the problem the most customers share, opens the pull request that fixes it, and measures the result.

1,284signals this week
9problems found
$184kat risk in the leading one
Signals per week

Live demo, illustrative data. Click a sheep to inspect it, or click the grass to lead the flock. Pick a week to replay the season.

Where signals graze. The tools you already pay for.

Connect in minutes. Personal data is redacted before anything is stored.

TicketsZendesk, IntercomErrorsSentry, OpenTelemetryCallsGongAnalyticsPostHog, AmplitudeReviewsApp Store, G2CommunitySlack, Discord

One loop, from complaint to measured fix. Then around again.

Here is one real problem going all the way through: CSV exports failing for 12 accounts.

BellwetherYou

  1. 01
    Ingest

    Tickets, errors, calls and reviews flow in. Personal data is redacted first.

    212 errors · 37 tickets · 4 calls
  2. 02
    Cluster

    Bellwether sees they are one problem and ties it to the accounts and revenue it touches.

    OPP-118 · 12 accounts · $184k ARR
  3. 03
    Spec

    It writes the change, the acceptance criteria and the number that will judge it.

    export_success_rate ≥ 95%
  4. 04
    Code

    Your coding agent builds it in a sandbox, behind a flag, and opens a pull request.

    PR #218 · 148 tests pass
  5. 05
    Ship

    You review and merge. Rollout is gradual, and guardrails roll back anything that regresses.

    1% → 100% · no guardrail tripped
  6. 06
    Measure

    The result is written back to the problem and the customers who asked are told.

    62% → 96% · verdict: win
  7. The verdict re-ranks the field, and the next bellwether leads.

The pens. Every problem has a place.

Problems move right as Bellwether and your team work on them. Open a card to see what is behind it.

Every pull request arrives with its reasons. No evidence, no PR.

Reviewers see who a change is for and how it will be judged before they read a line of code.

Every number is computed from stored evidence, not written by a model, so “12 accounts affected” is a query you can rerun.

Stream CSV exports instead of loading every row #218

MergedMaya merged 1 commit into main from fix/stream-csv-export
bellwether opened this for OPP-118

12 accounts ($184k ARR) cannot export tables over 10k rows. Exports load every row into memory and time out.

  • sentryTypeError in /export, 212 events in 7 days
  • zendesk“Export is greyed out for our big tables”
  • gong“We’d pay for a bigger plan if exports worked”
  • posthog31.4% of export_clicked never reach export_done
Judged by
export_success_rate, 62% → at least 95%
Guardrails
p95 API latency, error rate, auto-rollback
src/export/csv.ts+3−2
40export async function exportCSV(query, stream) {
41 const rows = await db.select(query)
42 return toCSV(rows)
41 for await (const page of db.pages(query, 1000)) {
42 stream.write(toCSV(page))
43 }
44}
  1. All checks passed · 148 tests, typecheck, guardrails armed
  2. Maya approved these changes
  3. Maya merged behind the flag export_streaming
  4. Rolled out 1% → 10% → 50% → 100%, no guardrail tripped
bellwether measured this change after 7 days
Win+34

export_success_rate, 62% → 96% after merge · dotted line is the 10% holdout

Autonomy is a dial. You choose how far it roams.

Set it per type of change. Copy fixes can ship on their own while new features still wait for a person.

Bellwether does itBellwether drafts, you approveYou do it

Authentication, billing, data migrations and anything touching personal data need a person at every level, by default.

Three bets. Why Bellwether is different.

PostHog, Amplitude and LogRocket each shipped an agent that turns signals into PRs in 2026. Each one runs mostly on its own data.

Any source

Works on top of the tools you already have, and combines what customers say with what they do.

ZendeskSentryGongPostHog+7 more
Your coding agent

Writing code is a commodity. Bellwether decides what to build and checks that it worked, and your agent writes it.

Claude CodeCodexCursorCopilot
Every change judged

No evidence, no problem. No success metric, no pull request. Every shipped change ends with a verdict.

WinFlatRolled back

Shepherd’s log. Every loop ends with a verdict.

Wins, flat results and rollbacks are all written down, and the ranking learns from each one. Illustrative data.

  1. Week 40
    Billing page timeoutPR #204 · cache plan lookups
    p95 6.1s → 0.9sWin
  2. Week 39
    Onboarding checklist copyPR #199 · shorter step titles
    activation 41% → 42%Flat
  3. Week 38
    Bulk invite modalPR #192 · new layout
    errors +3.1%Rolled back
  4. Week 37
    Password reset email delayPR #188 · move to a queue
    tickets −64%Win

Join the first flock. Make the next thing you ship the thing they asked for.

We are working with a small group of design partners who have real users, real logs, and more feedback than time.

Become a design partnerhello@bellwetherlabs.dev