How to Build a Growth Experiment Backlog

A growing tactic list creates more choices without making the next decision clearer. “Try outbound,” “publish content,” and “run ads” may all sound reasonable, but none tells a solo founder which uncertainty to address first.

To build a growth experiment backlog, start with the riskiest GTM assumptions behind one evaluated idea. Turn each assumption into a comparable bet with a hypothesis, channel, metric, and target. Score Impact, Confidence, and Ease consistently, sort the queue, check time, budget, access, and evidence constraints, then choose one next experiment.

This process starts after you have evaluated a SaaS idea before building. It ends when one bet is selected for the next evidence cycle. The backlog is a decision queue, not proof that its first item will work.

What a Growth Experiment Backlog Is

A growth experiment backlog is an ordered queue of comparable, testable bets for one selected idea. Every row uses the same minimum record and can be assessed against the founder’s current evidence and constraints.

The output is both a ranked set of candidates and a human decision about which one deserves attention next. That distinction matters. A score makes the review order explicit; it does not predict success or choose the experiment automatically.

A growth experiment backlog is not:

  • a brainstorm containing every possible channel
  • a marketing calendar that says when content or campaigns will ship
  • a product-development backlog containing features and technical tasks
  • an archive of completed experiments and results
  • a list of famous growth hacks borrowed from unrelated companies

It also answers a different question from ranking SaaS ideas before building. Idea ranking compares several possible businesses. An experiment backlog compares several evidence-generating bets for one business.

A tactic list is not an experiment backlog

Entries such as try outbound, do SEO, or post on LinkedIn name activities. They do not define what the founder expects to observe or what decision the result could change.

A backlog-ready candidate replaces the vague activity with a bounded proposition:

If we do X, we expect Y because Z.

For example, “try outbound” might become: “If we send a workflow-specific message to matched small-agency owners, we expect qualified replies because the message names a painful reporting task in their current workflow.”

The completed record still needs a metric, numeric target, channel, and comparative score. If no plausible result would change what the founder does next, the item is not ready for the ranked queue.

Step 1: Start With an Evaluated Idea's GTM Assumptions

Keep one idea per backlog. Mixing unrelated products creates incomparable Impact, Confidence, and Ease ratings because each idea has a different buyer, objective, access path, and cost structure.

Review the selected idea’s GTM approach, buyer evidence, and unresolved risks. Look for beliefs about the first reachable customer, channel access, message specificity, willingness to pay, activation friction, or manual delivery. Strategyzer similarly recommends starting an idea test with its most critical hypotheses.

Choose uncertainties whose answers could change a near-term decision. Do not start with a channel merely because it is popular.

Genhone’s empty experiment state can display the idea’s refined GTM approach as source context. The founder still identifies the assumptions and writes the candidates manually. Genhone does not extract assumptions or generate experiments from that context.

Extract assumptions, not channel ideas

The following is a planning worksheet, not a Genhone screenshot or stored product schema. All scenario details in this demonstration are illustrative.

Evaluated input Uncertain assumption Candidate test Decision the result could change
Small-agency owners appear reachable through an existing prospect list. The available list contains enough matched buyers to produce decision-relevant conversations. Send a narrow workflow message to matched agency owners. Continue with direct outreach or revisit the audience and access path.
The reporting and client-update workflow appears more urgent than the broad software category. Workflow-specific language will attract more qualified interest than category language. Test a narrow landing-page message with paid traffic. Keep, narrow, or replace the message before investing further.
Manual delivery appears feasible before product automation. Some agencies will commit to a paid reporting pilot without a finished product. Offer a manual reporting pilot to warm agency contacts. Pilot the service, change the offer, or revisit willingness to pay.

This worksheet can use context from Genhone or another documented evaluation process. Its source assumption and decision columns are working notes; Genhone does not store either as separate experiment fields.

Step 2: Turn Each Assumption Into a Backlog-Ready Bet

Convert each selected assumption into one record with a concrete title, a founder-written hypothesis, one supported channel, one observable metric, and a numeric target.

Make the audience and action specific. Prefer a metric close to the assumption: qualified replies for message and buyer-access uncertainty, paid-pilot commitments for willingness-to-pay uncertainty, or activation completions for onboarding uncertainty. Sending 100 emails is an activity count, not an outcome metric, unless delivery itself is the assumption being tested.

Set the target before starting. Targets in the demonstration below are illustrative decision thresholds, not conversion benchmarks or universal recommendations.

Use the minimum record

Field What the reader should write Product truth
Title A concrete test name, not only a channel. Required, 2–120 characters.
Hypothesis One bounded action, expected outcome, and reason. Required, up to 1,000 characters; founder-written.
Channel The route through which the test runs. One of the implemented channel labels.
Metric One observable measure close to the assumption. Required text field.
Target The numeric threshold declared before execution. Required numeric field.
Impact / Confidence / Ease Three comparative ratings using the same interpretation. Founder-entered integers from 1–10.

A Genhone record does not include separate fields for an objective, baseline, exposure, duration, owner, guardrail, score rationale, budget, dependency, source assumption, or branching decision. Those details can inform human judgment without being presented as product fields.

Apply a readiness gate before scoring

A candidate belongs in the ranked table only when:

  • it tests one identifiable assumption
  • its metric and numeric target can be declared before starting
  • its action is feasible with current audience access, time, and budget
  • at least one plausible result could change a near-term decision
  • it is not a duplicate of another candidate with different wording

Remove or rewrite anything that fails this gate. Scoring a vague tactic only gives the vagueness a decimal.

Step 3: Put the Candidates Into a Reusable Backlog Table

Copy this table into a document or spreadsheet. Rank is derived after scoring; it is not an independent input.

Rank Experiment title Hypothesis Channel Metric + target Impact Confidence Ease ICE
[Derived] [Concrete test] If we…, we expect… because… [Channel] [Metric]: [target] [1–10] [1–10] [1–10] [(I+C+E)/3]
1 Narrow workflow message to matched agency owners If we send a workflow-specific message to matched agency owners, we expect qualified replies because it names a current reporting pain. Cold Outreach Qualified replies: 4 8 7 8 7.7

The completed row is illustrative demonstration data, not a benchmark or Genhone result. The reusable structure does not require a product signup, downloadable file, or automated scoring tool.

Step 4: Prioritize With Lightweight ICE Scoring

Rate each qualified candidate on three dimensions:

  • Impact: how much a positive result could move the idea’s current growth objective or unblock an important decision.
  • Confidence: how strongly current observations, access, and prior evidence support the expected effect—not how certain the founder feels.
  • Ease: how feasible the test is for one founder given time, cost, dependencies, setup, and measurement effort.

Write the reason for each rating before choosing the number. Then apply the same interpretation to every candidate in the queue. Genhone does not infer, research, or generate these ratings.

Use Genhone's current formula consistently

Genhone uses this arithmetic:

ICE score = (Impact + Confidence + Ease) / 3

Each input is an integer from 1–10. The displayed score is rounded to one decimal, and backlog cards are sorted from highest to lowest score.

This is Genhone’s current formula, not a universal ICE definition. Current sources use different formulations: Growth Method presents an average, while ProductLift describes a multiplicative variant. The worked example below uses the Genhone average consistently.

A score of 7.7 is an ordering convenience. It means one candidate currently compares favorably with the others under the stated reasoning. It does not mean the experiment has a 77% success probability or that its future performance has been measured precisely.

Step 5: Rank the Queue and Choose One Next Experiment

Sort candidates by ICE, then review the leading rows against present constraints:

  • Would the result change a decision that matters now?
  • Can the founder reach the audience and observe the metric?
  • Does the test fit current time, cash, and implementation capacity?
  • Is a prerequisite or dependency missing?
  • Could another candidate produce useful evidence sooner?

Select one next bet and leave the other qualified candidates in the backlog. For a constrained solo founder, focusing on one active bet at a time can make interpretation and attention simpler. That is operating guidance, not a Genhone limit; the product can hold multiple running experiments.

The score starts the review. It does not finish it. A top-scoring interview campaign may depend on access to buyers the founder does not currently have. In that case, a slightly lower-scoring but reachable candidate can go first. Record the constraint rather than silently changing numbers to force the desired order. Rescore only when the evidence or effort estimate changes.

In Genhone, the founder selects Start to move a backlog item to running. Score order neither selects nor starts it automatically.

Use judgment when the score and constraint check disagree

A score should remain traceable to its reasons. If an apparent dependency was already reflected in Ease, do not penalize the candidate again during the constraint check. If the dependency is new, document it and update the relevant rating.

The final decision should be explainable in one sentence: “This bet goes next because it addresses the nearest important uncertainty with the best current combination of access, speed, and decision relevance.”

Stop there. Execution, result interpretation, verdicts, and retained learnings belong to the wider tracking process, not backlog selection.

Worked Example: Rank Four Experiments for One SaaS Idea

Demonstration: The idea, audience, actions, exposure details, targets, ratings, and scores below are illustrative. They are not Genhone results, conversion benchmarks, customer data, or recommended universal thresholds.

Assume a solo founder has evaluated a B2B SaaS idea for small agencies. The founder believes a narrow reporting and client-update pain may support a software product, but several GTM assumptions remain uncertain.

Show the source assumptions

The queue begins with three beliefs:

  • small-agency owners can be reached directly through an existing prospect list
  • a workflow-specific message will produce more decision-relevant replies than broad category language
  • some agencies may commit to a manual paid pilot before the product is automated

These assumptions and the resulting candidates were written for this demonstration. Genhone did not discover or generate them.

The exposure descriptions below make each action understandable. Exposure is surrounding planning context, not a separate Genhone field.

Show the ranked backlog

Rank Candidate Full illustrative hypothesis Channel Metric + illustrative target I / C / E Arithmetic ICE
1 Narrow workflow message to matched agency owners If we send a workflow-specific message to 30 matched small-agency owners from an existing prospect list, we expect 4 qualified replies because the message names a reporting and client-update pain in their current workflow. Cold Outreach Qualified replies: 4 8 / 7 / 8 (8 + 7 + 8) / 3 = 7.666… 7.7
2 Offer a manual reporting pilot to warm agencies If we offer a manually delivered reporting pilot to 10 warm agency contacts, we expect 2 paid-pilot commitments because immediate client-update relief may be valuable before the workflow is automated. Product-led Paid-pilot commitments: 2 9 / 5 / 4 (9 + 5 + 4) / 3 = 6.0 6.0
3 Publish a workflow-specific search page If we publish one search page for small-agency owners dealing with this reporting workflow, we expect 2 qualified discovery calls during the illustrative review period because the page addresses a specific problem they may actively research. SEO / Content Qualified discovery calls: 2 7 / 4 / 5 (7 + 4 + 5) / 3 = 5.333… 5.3
4 Send traffic to a narrow landing-page message If we send an illustrative fixed-budget campaign to a landing page about this narrow agency workflow, we expect 3 qualified demo requests because the segment- and pain-specific message should filter for more relevant interest than broad category copy. Paid Ads Qualified demo requests: 3 8 / 3 / 3 (8 + 3 + 3) / 3 = 4.666… 4.7

The rating rationales remain visible because the numbers alone cannot explain the order:

Candidate Impact rationale Confidence rationale Ease rationale
Narrow workflow message 8: Qualified replies would clarify buyer access and message relevance for the next GTM decision. 7: The founder has an existing matched list and a defined workflow pain, although neither guarantees replies. 8: One founder can prepare and send the bounded outreach with limited setup and direct measurement.
Manual reporting pilot 9: Paid commitments would provide a commercially relevant signal about the offer. 5: Warm access helps, but no supplied evidence shows that agencies will pay for manual delivery. 4: The offer, sales conversations, and manual fulfillment require more setup and capacity.
Workflow-specific search page 7: Qualified calls could improve understanding of the pain and organic acquisition path. 4: The scenario supplies no search-demand or current-ranking evidence for this page. 5: Publishing is feasible, but discovery, measurement, and signal arrival may be slower.
Narrow paid landing page 8: Qualified demo requests could clarify whether the message earns serious interest. 3: The audience, acquisition cost, and paid-message performance are currently unproven. 3: The test requires spend, tracking, landing-page setup, and campaign management.

The ranked order puts Narrow workflow message to matched agency owners first. It combines direct buyer access, fast setup, and a result that can change the immediate messaging and channel decision. The 7.7 does not predict success. It only makes the current comparison explicit.

The founder should select this Cold Outreach candidate as the next bet and retain the other qualified candidates in the queue. No result, verdict, learning, revenue outcome, or causal claim is implied by the demonstration.

How Genhone Records and Sorts the Backlog

Genhone experiment access is authenticated and attached to one idea in executing status. Before an idea has experiments, the empty state can display its refined GTM approach as source context. Once records exist, that source context is not stored as a separate field on every backlog card.

Creating a Genhone experiment requires a title, founder-written hypothesis, channel, metric, numeric target, and 1–10 Impact, Confidence, and Ease ratings. New records start in backlog.

The supported channel labels are SEO / Content, Social, Communities, Cold Outreach, Paid Ads, Partnerships, Product-led, Marketplaces / Directories, PR / Launch Platforms, and Other.

The interface displays the average ICE score rounded to one decimal and sorts backlog cards from highest to lowest score. A founder can edit, permanently delete, or start an item. Starting moves it to running.

Genhone does not provide drag-and-drop ordering, manual rank overrides, bulk import, tags, owners, approvals, dependencies, or review reminders.

State the product boundary plainly

Genhone’s tracker is manual. It does not:

  • generate growth strategy, assumptions, experiment ideas, hypotheses, or ICE ratings
  • automatically select or start the next experiment
  • execute experiments, send outreach, or buy traffic
  • ingest analytics or populate actual results
  • run A/B tests, assign audiences, manage feature flags, or calculate statistical significance
  • enforce only one running experiment
  • prove causality, demand, product-market fit, or future growth

The authenticated tracker is not a publicly inspectable tool, and no indexable tracker feature page currently exists.

Keep the Backlog Small and Current

There is no universal ideal backlog size. Keep enough ready candidates to compare, but not so many that the founder can no longer explain why each item exists or plausibly run it.

Revisit the queue when buyer evidence, channel access, cost, dependencies, or founder capacity changes. A mandatory weekly ceremony is unnecessary. Edit a score only when its underlying reason changes.

Remove duplicates and stale candidates instead of maintaining an idea cemetery. Genhone backlog items can be edited or permanently deleted; the current interface does not provide a separate parked status. Keep completed-result analysis outside this queue so the backlog remains focused on the next decision.

FAQ

What is a growth experiment backlog?

A growth experiment backlog is an ordered queue of comparable, testable bets for one selected idea. Unlike a tactic list, every item names an expected outcome, metric, target, and rationale. Unlike a product-development backlog, it organizes evidence-generating GTM tests rather than features or engineering tasks.

What fields should a growth experiment backlog include?

Use a concrete title, founder-written hypothesis, channel, metric, numeric target, Impact, Confidence, Ease, and the derived ICE score. Additional planning notes may help judgment, but they should not be confused with fields supported by a specific tracker.

How do you prioritize a growth experiment backlog?

Rate every ready candidate with the same interpretation of Impact, Confidence, and Ease. Genhone calculates (Impact + Confidence + Ease) / 3, rounds to one decimal, and sorts the backlog in descending order. Review the leading candidates against current access, measurement, capacity, cost, and decision relevance before choosing one.

Does the experiment with the highest ICE score always go first?

No. The highest score leads the review, but a missing dependency, unreachable audience, weak measurement path, or more urgent strategic question can change the next choice. Keep the scoring rationale visible, document the constraint, and rescore only if the underlying evidence or effort estimate changes.

How many experiments should be in a backlog?

There is no universal number. Keep enough qualified candidates to make a useful comparison without retaining items the founder cannot explain, measure, or plausibly run. Remove duplicates and stale bets as evidence and constraints change.

Can I keep a growth experiment backlog in a spreadsheet?

Yes. The copyable table in this article is deliberately tool-agnostic. A dedicated tracker becomes useful when attachment to one evaluated idea, consistent fields, score-based ordering, and retained context matter, but the core method does not depend on a particular product.

Turn One GTM Assumption Into the Next Bet

A useful growth experiment backlog turns evaluated uncertainty into comparable candidates, sorts them transparently, and ends with one deliberate choice.

Start with the assumption whose answer could change the nearest real decision. Write the bet, declare the metric and target, record the reasons behind the scores, then apply the final constraint check. Choose the experiment that best combines decision relevance with present access and capacity—not the channel that happens to be fashionable.

About the author

Malte Hedderich is a machine learning engineer and the founder of Genhone. He works on AI, MLOps, and agentic software workflows, and writes about machine learning and AI systems at hedderich.pro.

  • Machine learning engineer with experience in artificial intelligence and MLOps.
  • Master of Science in Business Informatics from the Technical University of Darmstadt; studied Software Engineering at Tongji University in Shanghai.
  • Has shipped multiple SaaS or software products and uses LLM-powered and agentic coding workflows.
  • Has firsthand experience with the build-before-validation failure pattern.