Bug Buster · Operational control
A bug report is evidence, not permission to change production
AI can classify reports, review pull requests, correct lint failures, reproduce faults, and prepare fixes. Bug Buster escalates regressions and risks introduced by a change instead of treating a clean diff as permission to release.
The operating problem
The queue is not slow because nobody can type the fix
A production report usually begins incomplete: “booking failed”, “the total looks wrong”, or “I cannot sign in”. The costly work is turning that observation into a bounded claim. Which user state produced it? Is it repeatable? Did data become incorrect or did the interface only display it incorrectly? Which products share the dependency? Is there already an incident, a duplicate, or a known limitation?
AI can reduce the search cost across reports, logs, traces, tests, and code. That makes it valuable early in the process, before a line is changed. It can assemble context and test a hypothesis quickly. What it cannot infer from a ticket is the authority to decide that the proposed behaviour is correct for the organisation.
Maintenance accelerates when evidence moves faster—not when production authority becomes ambiguous.
Decision path
A signal earns investigation before it earns a change
01 · decision
Report or PR signal
Capture the observed outcome, failed check, changed lines, and relevant state.
02 · decision
Reproduce or verify
Create a failing example, rerun the exact rule, or mark the uncertainty explicitly.
03 · decision
Bound impact
Identify affected systems, data, authority, and reversibility.
04 · evidence
Prepare evidence
A proposed change carries tests, scope, and a rollback path.
05 · refused
Release authority
The product owner decides whether the evidence is sufficient for this risk.
The plugin model
Bring product context to one triage discipline
Bug Buster operates as a common maintenance workflow paired with a lightweight product integration installed as aversioned dependency in each system (such as Garage CRM or Grand Total). The product-side plugin supplies bounded context: release identity, feature area, safe diagnostics, ownership, test commands, and links to the applicable runbook. The central service groups related reports, proposes severity, traces the affected surface, and prepares a reviewable change where the evidence is sufficient.
On an open pull request, the integration can run approved lint and test commands, connect failures to changed lines, and prepare mechanical corrections. New invariants, error paths, permission changes, migration hazards, and dependency risks stay attached to the pull request and are escalated to its reviewers rather than hidden inside a broader fix.
The plugin is not a permanent administrative tunnel. It should expose the least authority required for observation and preparation, with explicit capability scopes and an auditable request trail. Secrets, personal data, and unrestricted production access do not become acceptable merely because the consumer is an internal automation.
| Stage | Automation may | Product owner retains |
|---|---|---|
| Intake | Deduplicate reports, inspect pull-request signals, and request missing context | Definition of material impact and escalation obligations |
| PR preflight | Run approved lint, test, and static-analysis commands; prepare mechanical corrections | Decision on findings introduced by the change and whether the pull request may proceed |
| Reproduction | Build a failing example in an isolated environment | Judgement about whether the example reflects intended behaviour |
| Prioritisation | Estimate scope, affected versions, and evidence confidence | Trade-off between user harm, operational risk, and planned work |
| Repair | Prepare a small change, tests, release notes, and rollback evidence | Approval of behaviour, migration, and release timing |
| Learning | Cluster recurring fault patterns and propose guardrails | Decision to change architecture, policy, or product design |
Pull-request guard
Fix the lint error. Escalate the behaviour change.
Deterministic lint failures such as formatting, import order, or an unused symbol can usually be corrected and verified by rerunning the same rule. A failing test may also have a narrow correction when the intended behaviour is already explicit.
A pull request can expose a larger issue: an omitted authority check, a retry that duplicates an external action, an incompatible migration, or a test changed to accept the wrong domain behaviour. Bug Buster surfaces the introduced risk, identifies the affected boundary, and stops the automated repair path. Rewriting it silently would remove evidence the reviewers need.
| Pull-request finding | Bug Buster response | Escalation condition |
|---|---|---|
| Deterministic lint failure | Prepare the smallest correction and rerun the exact lint rule | The correction changes runtime behaviour or crosses files outside the declared scope |
| Regression test failure | Compare the base and proposed revisions; fix only when the intended rule is explicit | The pull request changes the rule, test data, or expected outcome |
| New security or authority path | Describe the changed boundary and attach the relevant evidence | Always require an accountable reviewer; never auto-approve |
| Migration or compatibility risk | Identify affected versions and the missing forward or rollback path | Block until adoption order and recovery are decided |
| Unrelated existing defect | Create a separate finding linked to the evidence | Do not enlarge the current pull request merely because a nearby defect was found |
The cleanest pull request is not the one with no warnings. It is the one whose remaining decisions are visible to the people authorised to make them.
Autonomy ladder
Widen evidence before widening authority
01 · decision
Observe
Summarise and classify reports without code access.
02 · decision
Recommend
Suggest priority, affected area, and reproduction steps.
03 · evidence
Prepare
Create a bounded change with tests, but never publish it.
04 · refused
Release narrowly
Only a defined, reversible fault class may cross the gate automatically.
The last step is not the target state for every team. Most value arrives while release authority remains human.
Authority
Autonomy should follow fault classes, not model confidence
A model confidence score is not a production control. It describes a system's internal assessment, not the consequence of being wrong. A high-confidence change to tax calculation may deserve more scrutiny than a lower-confidence correction to an internal label. The useful unit of authority is therefore a named fault class with known blast radius, test evidence, reversibility, and an accountable owner.
An organisation might permit automatic release of a deterministic documentation correction or a dependency patch that passes a defined compatibility suite and can be rolled back without data change. It should not generalise that permission to “high-confidence bugs”. Each additional class is earned through observed performance and incident review. We examined the operational boundaries, safeguards, and interactive refusal queue for this model in Bugs fixed overnight, reviewed in the morning.
- Start with read-only observation and make the evidence useful before permitting code preparation.
- Require a reproducible failure or explicitly label the change as a hypothesis.
- Constrain change size, affected paths, dependency scope, and permitted test commands.
- Treat schema, identity, money, and irreversible external actions as separate high-consequence classes.
- Revoke authority automatically when evidence is missing, telemetry is degraded, or rollback is unavailable.
Reproduction
A reproducible failure is a maintenance artifact, not a comment on a ticket
“Confirmed” is too weak a state for an automated maintenance path. A useful reproduction names the starting state, the action, the observed result, and the result that should have occurred. It identifies the product and release, controls time and external dependencies where possible, and fails for the same reason as the reported problem. Without that last condition, a passing fix may only silence a test that never represented the incident.
The best form is an automated test at the lowest boundary that still preserves the fault. A calculation error may need a unit test; a duplicate reservation may require a concurrency test against a real database; a provider-specific sign-in failure may need a contract fixture and a browser path. For an intermittent production fault, the first artifact may instead be a trace, a sanitised event sequence, or a deterministic replay. The format follows the failure rather than a universal testing preference.
| Reproduction quality | What it establishes | What remains unknown |
|---|---|---|
| Report only | A user observed an unwanted outcome | Starting state, frequency, affected versions, and cause |
| Repeated manually | The outcome can be produced under known steps | Whether the repair will remain protected after release |
| Automated at the boundary | The failure is repeatable and can become a regression gate | Whether adjacent workflows or historic data also need repair |
| Production trace or replay | The real sequence and system state are represented | Whether sensitive data is safe to retain and whether replay changes external systems |
When reproduction is impossible, Bug Buster should preserve that uncertainty. A change can still be prepared as an instrumented hypothesis: add a narrow diagnostic, improve a refusal signal, or guard a suspected state transition. It should not be presented as a verified fix. Keeping “observed”, “reproduced”, and “explained” as separate states prevents the queue from turning confidence into fact merely because a plausible patch is available.
Prioritisation
Severity is a decision about consequence, not message volume
Ten identical reports may describe one contained inconvenience. One quiet reconciliation failure may corrupt every future statement. Prioritisation needs several dimensions: current user harm, data integrity, security or authority impact, number and identity of affected parties, persistence, detectability, workaround quality, and reversibility. Frequency matters, but it does not stand alone.
| Dimension | Question | Escalating evidence |
|---|---|---|
| Authority | Can a person see or do something they should not? | Cross-tenant access, privilege persistence, missing revocation |
| Integrity | Can the system create an incorrect durable state? | Financial mismatch, duplicate action, unreconciled record |
| Reach | Who is affected now and on which versions? | Growing cohort, shared dependency, no safe containment |
| Recovery | Can the outcome be reversed without losing legitimate work? | Manual repair, external side effect, uncertain history |
| Visibility | Would ordinary monitoring reveal continued harm? | Silent drift, successful responses, delayed discovery |
Operations
Measure whether the maintenance system improves control
Closing more tickets can conceal worse maintenance. Useful operating signals include time to a reproducible example, the age of unbounded high-consequence reports, escaped regressions, reopened changes, rollback frequency, and the proportion of prepared fixes rejected because the intended behaviour was wrong. Those rejections are not wasted automation; they show the review boundary is doing work.
Every prepared change should retain its lineage: original observations, gathered diagnostics, reproduction, affected versions, code change, test evidence, approval, release, and post-release result. Incident review can then improve both the product and the common maintenance path. A recurring identity defect may become a shared contract test; a repeated pipeline escape may become a platform gate.
The system is learning when one product's failure becomes a guardrail for the rest of the ecosystem.
The counter-case
Some reports should never enter an automated repair path
Security disclosures, suspected fraud, legal holds, sensitive personnel matters, and active safety incidents need deliberately limited handling. They may require a protected intake, restricted evidence access, preservation rules, and coordination that ordinary triage must not automate or expose. The correct plugin behaviour is to recognise the class, capture the minimum, and transfer authority.
Small products with low change volume may also need only a disciplined checklist and existing delivery pipeline. Bug Buster earns its shared role when products repeat the same evidence-gathering and maintenance controls. It should not manufacture process where the coordination cost exceeds the risk it removes.
Let's talk about your challenge
If your organization is working with complex digital systems or exploring operational AI, we are always open to a conversation.