Phrom: AI-Assisted Backlog Quality Checks
Designing and building a human-in-the-loop tool that reviews GitHub Issues before backlog refinement, combining versioned deterministic rules with local LLM-based suggestions while keeping every product decision with the Product Owner.
Snapshot
- Role: Product Owner, Solution Designer and Developer (Solo Project)
- Status: MVP, Version 1.0.0
- Period: October 2026
- Stack: Node.js CLI, GitHub REST API, Ollama (
qwen3:30b-instruct), JSON-based rules engine
Summary
Phrom checks GitHub Issues before backlog refinement against a transparent, versioned ruleset. It shows which items are ready and generates a concrete improvement draft for all others. The AI runs locally via Ollama, and Phrom never modifies issues itself: the Product Owner reviews every suggestion and applies changes manually.
The case study describes the problem, the architecture decisions, the MVP outcome and the lessons learned from building it. Effectiveness and model quality are not yet measured; more on that in the section βWhat remains openβ.
1. Starting Point and Problem
During refinement, the team spends time on work that could be done beforehand: unclear story formulations, missing or non-testable acceptance criteria, oversized items and duplicates. The conversation then starts with wording questions instead of the discussion that really needs team time: jointly understanding, estimating and choosing the technical solution path.
Phrom deliberately addresses only one part of this problem, namely poorly prepared items. Prioritisation, estimation, feasibility and sprint readiness remain explicitly with the team.
Target audience: Product Owners who maintain a backlog in GitHub Issues, need to sharpen items before every refinement, have no dedicated tooling budget and do not want to send their backlog content to an external LLM provider.
2. Goals and Guardrails
| Guardrail | Consequence in design |
|---|---|
| AI must not decide | Human-in-the-loop: Phrom never writes to GitHub, the PO takes over manually |
| Assessment must be traceable | Ruleset in criteria/*.json, each criterion marked as (deterministic) or (ai) |
| Data stays under control | Local model via Ollama, no external LLM provider |
| Minimal permissions | Fine-grained token, single repository, for analysis only Issues: Read-only |
| No magic | Fixed pipeline instead of agent, no framework, no model tuning |
3. Solution
Flow
The PO starts the CLI, Phrom reads the issues via the GitHub REST API and evaluates each issue in two layers:
- Score (0β100): points achieved relative to points achievable.
- Ready Gate: All criteria marked as
requiredmust pass.
| Status | Condition |
|---|---|
| π’ Ready | Score β₯ 80 and Ready Gate passed |
| π‘ Needs work | Score 50β79 |
| π΄ Not ready | Score < 50 or Ready Gate not passed |
The Ready Gate overrides the score. An issue with 73 points and non-testable acceptance criteria (required) is therefore π΄ Not ready.
Two Types of Criteria
- Deterministic (rule): checked in code, in milliseconds, reproducible, e.g. story format, epic link or number of acceptance criteria.
- AI: the local model provides a reasoned assessment of semantic questions, e.g. whether acceptance criteria are measurable. This is explicitly an assessment, not a fact.
The issue type (story, epic, task, bug) determines which criteria apply. Points and required markers can be adjusted per team in criteria/*.json.
Architecture
PO
β
βΌ
Node.js CLI (phrom)
β
βββ GitHub REST API ββββββββββββββΊ GitHub Issues
βββ Rules Engine βββββββββββββββββΊ criteria/*.json
β βββ Deterministic Checks ββββΊ checks.js
β βββ Ready Gate βββββββββββββββΊ criteria-loader.js
βββ AI Engine ββββββββββββββββββββΊ model.js β Ollama
βββ Improvement Engine βββββββββββΊ improve.js β model-improve.js
β βββ Reference Templates ββββββΊ references/*.json
βββ Filesystem βββββββββββββββββββΊ output/
β
βΌ
PO Review β Manual Adoption β GitHub Issue
Improvement drafts are generated via one-shot prompting: per issue type, the model receives a reference example from references/*.json. No model weights are trained. Where information is missing (role, benefit, epic number), Phrom does not invent anything, but leaves placeholders in square brackets.
4. Example: Issue #3 βImprove loginβ
A typical one-sentence story from the demo repository, without story formulation, context or epic link and with only one vague acceptance criterion.
Assessment: π΄ Not ready Β· 27/100, six gaps.
β Story-format Rule No story formulation found
β Story-context Rule Neither product nor target audience mentioned
β Epic-link Rule No epic reference
β AC-presence Rule Only 1 acceptance criterion (expected: happy path + error case)
β AC-testability AI Criteria vague and not measurable
β Business-value AI No concrete user, no measurable benefit recognisable
Excerpt from the generated draft (to be reviewed by the PO):
## Story
As a [role], I want to [specific login improvement], so that [concrete benefit], measurable by [metric].
## Acceptance Criteria
**Happy Path**
- [ ] Given I am on the login page and have valid credentials, when I enter my username
and password and click "Sign In", then I am redirected to the dashboard within 2 seconds.
**Error Cases**
- [ ] Given I enter an incorrect password, when I click "Sign In", then I see a clear error
message stating "Invalid credentials. Please try again." and the login form remains visible.
The placeholders are intentional: only the PO knows the role, benefit and metric.
5. Architecture Decisions
Two decisions are documented as ADRs.
ADR-001: GitHub Issues as data source. Compared were GitHub Issues, Trello and a self-hosted Kanban (Wekan). GitHub won because of the token that can be limited to one repository, the well-documented REST API, the portfolio effect (code, demo and results in one place) and because no additional infrastructure is needed. Known downsides: external dependency, unstructured free-text fields and no native Kanban via the REST API.
ADR-002: Two repositories for code and demo backlog. phrom contains code, rules and tests, phrom-backlog-demo exclusively demo issues. Compared were one repository and one repository with label filtering. The separation yields a clean scope, a token that only touches the demo repo, a safe reset and a clear presentation for interested parties. The price is more administration and configuration.
6. MVP Outcome
Implemented and runnable are:
- six CLI commands (
run,improve,select,filter,status,list), - a versioned ruleset (Version 0.3.0) for four issue types,
- improvement drafts with before-and-after view,
- reports as Markdown and JSON in the
output/folder, - a public demo backlog, usable via
npm run demowithout write permissions.
There are no measured statements on time saved or hit rate. Figures on this would not be substantiated at this point.
7. Lessons Learned from Building It
- Pipeline instead of agent. Phrom executes fixed steps, the model does not choose tools. This makes the behaviour reproducible and error search simple. An agent with tool selection, state management and resumption is a deliberate later expansion.
- Rules first, AI afterwards. Everything that can be checked deterministically runs without a model. The AI is reserved for semantic questions and is marked as such in every report.
- Demo reset has limits. The GitHub REST API does not know deleting issues, that only works via the GraphQL mutation
deleteIssuewith admin rights. In addition, the issue number counter of a repository cannot be reset: after deleting 12 issues, the next seed started at #13. Reports and tests should therefore map via stable identifiers instead of issue numbers. - Separate permissions. Analysis only needs read access. Write and delete rights belong exclusively to seed and reset, which run separately.
8. What Remains Open
Known limitations: no status tracking per issue, no resumption and no retry on crashes, static references, configuration only via JSON, no UI.
Roadmap:
- Quality measurement against a manually assessed test set and comparison of different models. The hypothesis that a local model is sufficient for the content checks is so far unproven.
- Persistence in PostgreSQL instead of files.
- Reference matching via similarity search instead of static reference JSONs.
- Status tracking, retry and resume for long runs.
- PO review UI and PO config UI.
- Optional write-back to GitHub, only after explicit approval.
- Evaluation of issue comments.
Links
- Code: github.com/MKalder/phrom
- Demo backlog: github.com/MKalder/phrom-backlog-demo