Todd Watts

Build a Software Roadmap Before Committing to a Rewrite

Todd WattsUpdated 10 min read

An original fictional exercise: inventory a service firm's software, compare competing priorities, and define a reversible first increment with explicit acceptance and rollback criteria.

A folded drafting plan links existing buildings, with a small removable lime bridge representing one bounded software change.
A smaller intervention can test the direction.Illustration · Todd Watts
04 / Reversible first increment

Add a view before moving the authority.

  1. 01 / Authoritative register

    Keep existing writes

    Job state stays in the register. Billing owns invoices. Approved reminders keep their existing owner.

    Dedicated read credentials only
  2. 02 / Candidate snapshot

    Validate, then replace

    Check IDs, statuses and completeness. Publish the snapshot as a whole; retain the previous one if refresh fails.

    No partial snapshot becomes current
  3. 03 / Coordinator queue

    Return permitted rows

    The server filters access before responding. The view shows freshness and links back for any action.

    No edits, invoices or reminders
Inspect the stop and rollback boundary

Stop on unauthorized data, an unexpected write or an unexplained missing job. A snapshot older than ten minutes is unavailable for planning.

Remove the queue entry point, disable its refresh job and revoke its read credentials. Confirm the original workflow still operates. This rollback does not cover a later migration that changes authoritative records.

Proposed design for the article’s fictional equipment-inspection firm. Gates and limits describe acceptance criteria, not completed tests or service levels.

Make the decision smaller than the entire system

Before approving a rewrite, I want a clear account of what fails, who depends on the current behavior and which change would test the most important assumption. A roadmap should connect those facts to a sequence of decisions. A list of replacement technologies does not do that on its own.

The exercise below is entirely fictional. Its service firm, software inventory, workload and test data were invented for this article. The proposed checks have not been run against a client system, and the numbers are acceptance criteria, not reported results.

The example firm schedules equipment inspections. Its owner wants a new customer portal because staff struggle to explain where jobs stand. Operations wants fewer missed handoffs. The developers want to replace an application they find difficult to change. All three concerns deserve attention, but they do not imply the same first project.

Inventory the responsibilities before the technologies

Start with the record each system owns and the people who rely on it. In this exercise, the inventory looks like this:

Scroll horizontally if needed to read every column.

Existing componentRecord or responsibility it ownsDependency to preserve
Request formThe original request and its stable request IDCoordinators must be able to trace a job back to its intake
Job registerCurrent job status and assigned coordinatorThe daily scheduling process reads and updates this register
Billing productInvoices and payment statusStaff must not infer payment from a job's operational status
Reminder workerSending a reminder after a coordinator approves itA refreshed status display must not trigger a second reminder
Shared daily reportA manually assembled view of open jobsOperations uses it to plan the next day's work

The assumed problem is narrower than “the application is old.” Staff assemble the daily report from several screens, and they cannot reliably distinguish a waiting job from one ready for scheduling. The first question is whether a clearer operational view would address that problem without replacing the systems that own the records.

This inventory also exposes unknowns. Are request IDs stable across exports? Who may see each job? Can the register supply a complete snapshot? Until those questions have answers, drawing a new interface is premature.

Compare the competing priorities

The following ordering is my proposal for these invented constraints. It is not a universal priority score.

Scroll horizontally if needed to read every column.

CandidateWhy someone wants itWhat must be learned firstProposed disposition
Customer portalCustomers could check their own progressWhich statuses are accurate, understandable and appropriate to expose?Defer customer access until the status model is established
Full application rewriteDevelopers want a system that is easier to changeWhich existing behaviors and dependencies must survive?Keep as an option; do not authorize from age or frustration alone
Automatic remindersOperations wants fewer manual follow-upsWhich event authorizes a reminder, and how are duplicate sends prevented?Defer outbound automation while status ownership remains unclear
Internal read-only queueCoordinators need a dependable view of work waiting or readyCan the current records support an accurate, permission-aware view?Investigate as the first bounded increment

The queue wins this comparison because it tests the shared dependency behind the other proposals: whether the business can describe a job's state accurately. It does not win because dashboards are inherently valuable. If an existing report can answer the same questions with a small change, improve that report instead.

Write down the decision and what could overturn it

An architectural decision record keeps the context, choice and consequences together. AWS describes this role for ADRs, including retaining the history when a later decision supersedes an earlier one. The compact record below is an original example, not an AWS template or a record of completed work. AWS guidance on architectural decision records

Scroll horizontally if needed to read every column.

FieldProposed decision for the fictional firm
StatusProposed, pending source-data and access checks
ChoiceTest an internal, read-only work queue before authorizing a wider rewrite
Keep authoritativeThe existing register owns job status; billing owns invoices; the reminder worker owns approved sends
First usersA small coordinator pilot with explicitly agreed access
First scopeShow ready, waiting and closed jobs with source IDs, freshness and unresolved-record warnings
ExcludedEditing job state, taking payments, customer access, sending reminders and migrating historical records
Main costAnother view to maintain, plus mapping and freshness checks
Revisit whenSource records cannot express the required states, access cannot be enforced, or the queue fails to answer the operational question

The proposed view cannot repair a poor status model by displaying it more attractively. A failed source-data check may lead to a small change in the existing register before any new interface is built.

Define a reversible first increment

For this exercise, the queue reads an authorized projection from the register. It has no write credentials to the register, billing product or reminder worker. It links coordinators back to the current application for any action that changes a job.

The projection contains only fields needed for the queue: job ID, request ID, operational status, coordinator and the time of the last completed refresh. Access is checked on the server before returning records. The interface does not download all jobs and rely on hiding unauthorized rows.

Use a complete snapshot for this first design. Build and validate the next snapshot separately, then replace the current snapshot as a whole. If retrieval stops halfway through, retain the last complete snapshot and show that refresh failed. A partial response must not silently make jobs disappear from the queue.

Assume a refresh is attempted every five minutes. For the pilot, a snapshot older than ten minutes is considered too stale for planning. The interface then replaces the actionable queue with a stale-data message and a link to the existing workflow. These are proposed limits chosen for the example, not measured service levels.

Because this increment cannot change the authoritative records or send messages, turning it off leaves those operations in the existing system. That makes this particular rollback small. It does not prove that a later migration with writes, new states or altered data formats would be equally reversible.

Make the acceptance checks specific

Create a synthetic fixture with 60 jobs: 20 ready, 25 waiting and 15 closed. Give one fixture user access to all 60 jobs, then assign restricted users to known subsets. Keep an explicit expected list of IDs and status counts for each user's view; matching a total alone could conceal both a missing job and an extra one.

Scroll horizontally if needed to read every column.

CheckRequired behavior before the pilot can proceed
Normal snapshotEach user receives exactly their authorized IDs and statuses. Only the full-access fixture user has totals of 20 ready, 25 waiting and 15 closed
Identical repeated source rowA repeated row with the same ID and values does not produce a second job
Conflicting rows for one IDPut that job in a visible “needs review” state; do not guess which status is correct
Partial source responseDo not publish the candidate snapshot; preserve the previous complete one and report the failed refresh
Unknown status valueFlag the record for review rather than silently treating it as ready or closed
Permission boundaryAn unauthorized job is absent from the server response, not just from the visible table
Expired snapshotAt more than ten minutes old, the queue cannot present itself as current; the existing workflow remains reachable
Reload or repeated viewNo job mutation, invoice action or reminder send occurs

The conflict fixture should deliberately change one of the 20 ready jobs. For the full-access fixture user, its expected state becomes 19 ready, 25 waiting, 15 closed and one needing review. A restricted user's counts change only if that job is in their authorized subset. This makes the treatment of uncertainty inspectable rather than hiding it inside a success count.

After the deterministic checks, ask pilot coordinators to perform a defined task: identify which jobs can be scheduled and explain why each remaining job is waiting. Compare their answer with the source records. Record errors and unresolved meanings before treating a faster interaction as an improvement.

Decide when to stop, roll back or continue

Stop the pilot immediately if a user receives a record they are not allowed to access, a read action changes an authoritative record, or a job disappears without an explicit explanation. Disable the queue and return to the existing workflow while investigating. Do not wait for a trend before responding to a broken boundary.

For a freshness failure, mark the view unavailable for planning and direct users to the existing application. Repeated freshness failures are a reason to reconsider the projection design or the chosen limit, not to remove the warning.

The rollback procedure for this first increment is to remove the queue entry point, disable its refresh job and revoke its dedicated read credentials. Confirm that the existing register and daily process still operate. Keep the decision and failure evidence so the next attempt does not rediscover the same problem.

If the pilot succeeds, decide the next responsibility to change. A useful queue might be the entire needed improvement. It might also provide evidence for replacing one bounded part of the application. Incremental replacement has its own costs: the old and new paths must coexist, and shared data responsibilities remain explicit. Microsoft's Strangler Fig guidance describes that tradeoff; this read-only queue is a preparatory experiment, not a complete implementation of the migration pattern. Microsoft's incremental migration guidance

What the roadmap has earned

The proposal now has a specific business question, an inventory of ownership, a first increment and conditions for stopping. It has not earned a claim about reduced cost, faster delivery or successful production migration. Those require evidence from actual implementation and use.

You can inspect a different bounded problem in the original lending lab, where duplicate requests and failed notifications are visible. The public project record separates independent technical work from client context.

If you need help connecting a business decision to architecture and implementation, fractional CTO support describes that work. For engineering hiring, my résumé provides the career and technical background behind these decision exercises.

Todd Watts

Software engineer and fractional CTO through Shell Command, LLC. Technical direction, architecture and hands-on development.

← All posts