GTM OS · GTM simulator

A GTM simulator that makes judgment calls.

A world model of your go-to-market. Aerospace simulates before it flies. Drug discovery simulates cells before it touches a patient. Go-to-market still tests in production, on your budget, your reputation and your buyers' patience. The simulator is where you run the experiment first. It tells you what to run, what to kill and when you will know.

300
buyers with memories
50
plays run forward
24
runs per plan
16
weeks per run
In plain words

It tells you what is going to work next.

Ten minutes of setup, then you can try any plan you like before you spend a euro on it.

IN

You describe your company

What you spend in a week. How many deals your team can work at once. How many of your buyers have heard of you. What you already run today. Eight numbers and a list.

RUN

It plays the next four months

Three hundred buyers, one week at a time, twenty-four times over so you see a range instead of a single guess. Then it does the same for every plan you want to try, on the same buyers with the same luck, so the plan is the only thing that changed.

OUT

You get a ranking and a reason

Which plan came out ahead most often, by how much, the week you would be able to tell in real life, plus what stopped the ones that failed. It ranks your options. Treat the euro amounts as illustrative until you feed it your own history.

The questions

Four questions every GTM leader asks. Three more they should.

01

What is going to work?

Your plans, ranked by how often each one finished ahead of doing nothing. Ahead in 70 runs out of 100 or better and it says works. Between 55 and 70 it says coin flip, because that is what it is.

02

What is not going to work?

The same list read from the bottom, at 55 out of 100 or worse, with the thing that killed each plan named out loud. You ran out of salespeople. You ran out of budget. Too many people had to say yes. The price was wrong. The buyers got tired of hearing from you.

03

What is my next best step?

The one play worth doing now, because it gains you the most or because it teaches you the most. You pick what to try and it shows you where that lands against everything else the same hours could have bought. Sometimes the honest answer is to stop doing something you already do.

04

What do I do after that?

Name the number you want to hit and the most you will spend in a week. It works backwards to find the order of moves that gets you there, or it tells you the target is out of reach on that money. Most teams start here, because they wake up wanting a number to move.

The three that rarely get asked

05

When will I know?

The first week the early signs pull away from doing nothing and stay there. Write that week down before you start, then stop if it comes and goes.

06

What would I have to believe?

Which single assumption, if it turns out to be 30 percent off, changes the whole ranking. That is the number worth going and measuring first.

07

How much should I trust this?

Paste your real wins week by week and it grades how close it gets to your own history. Tell it what a play actually did and it adjusts what it believes about that play. Until you do one of those, trust the ranking and the reasons behind it. The euro amounts are illustrative.

The split

Same buyers. Same dice.

The simulator calls each copy of your company an arm, the way a drug trial does. Everything below is two arms.

Three hundred buyers, drawn once. Both copies of the company meet the same three hundred, in the same order, with the same luck. Whatever separates them at the end came from the plays.

Arm A · outbound push0 won
Arm B · fix conversion0 won
Cumulative wins
0481216
The gap at week 16
+0

buyers won in A only

The gap is the change. Nothing else was different.

week 0
won won in A, missed here Illustrative shapes, drawn so A wins here. In the real run either arm can, on your numbers.
Try it now

Move a slider. Watch the answer change.

Simulatorillustrative mid-market · sales-led · 16 weeksRuns in your browser
Loading the run

It has already worked out what is holding this company back, picked the three plans most likely to fix it and run them. Change anything and it runs again.

How it works

Five steps, one loop.

01

State

Weekly budget, how many deals your team can work at once, outbound and marketing hours, how well known you are, how much you are trusted, how tight buyer budgets are, how hard competitors push. Those sliders are your company as it stands.

02

The plans being compared

Make copies and change one thing in each. Copy A goes hard on outbound. Copy B fixes conversion. Both start from the same day, so any difference at the end came from the change.

03

Play them forward

Each copy runs week by week. Reps fill up. Lists run down. Deals take their real time to close. Plays that do not fit the weekly budget pause that week.

04

The readout

Every plan is scored against doing nothing, because a plan only counts if it beats that. You get a direction, a range and a plain answer.

05

Log reality

Tell it what really happened when you ran a play. It adjusts what it believes about that play and the next run reorders. What this browser learns stays in this browser.

Why arms
In a drug trial one group gets the drug, one gets a placebo, everything else is kept identical. Same idea here, with your company in place of patients. The pill is a budget decision.
What it is

One living picture, built from four layers.

A picture of your go-to-market that plans can be tested against.

L1

Your state

Eight sliders. Weekly budget, deals worked at once, outbound hours, marketing hours, brand awareness, baseline trust, buyer budgets, competitor pressure. Plus the plays you already run.

L2

The buyers

Three hundred of them, each with a memory of every touch you send and a weight for how much it mattered. A peer referral lands heavier than an ad. Enough noise and they get tired, then they opt out.

L3

The play library

Fifty plays, each with a weekly cost, what it eats from your shared hours, how long before anything shows, plus a starting assumption about how much it helps. Every one runs forward in time.

L4

The learning loop

Tell it what a play really did and it adjusts what it believes about that play, so the next run reorders on your evidence. Paste your own weekly wins to see how far it sits from your history.

The part no spreadsheet does

Ask a simulated buyer why.

Pick one of the buyers you lost and ask. The answer is composed from that buyer's own memory, so it names the touches, the weeks and the reason.

Interviewsame company as the run above · lost dealsNo language model
Simulating so there is someone to interview
The clock

Why time changes the answer.

A list of good ideas puts them in order. It cannot tell you the third one breaks the second, or that the first has no chance of showing up inside the quarter you judge it in.

01

Seller capacity

A team can only hold so many first meetings a week. Anything booked past that ceiling sits waiting and goes cold.

02

List saturation

The further you work through your list, the worse each email does. Sending twice as much buys nowhere near twice the replies.

03

Pipeline lag

Most of what this quarter builds closes next quarter.

04

Compounding content

Pages take months to warm up. The same play reads flat over 90 days and strong over a year. That is the calendar talking. The play is fine.

05

Shared pools

Hours, budget and the list are all finite. Every play eats from the same pools, so two email plays on one list quietly starve each other.

06

Tired buyers

Keep touching the same people and they stop noticing, then they opt out. A plan that looks strong on volume can lose on how people feel about it.

07

Payback against the clock

Every plan prints what a customer cost you and how many months of margin it takes to pay that back, so a move that works and pays back too late is visible as that.

10weeks

Median mid-market deal in the run above. Most of what this quarter builds closes next quarter.

Under the hood

What is in there and what is not.

The simulator is the engine. The GTM OS is the harness around it.

What it is built from

The rules inside are written-down practitioner experience, the play library above. It is a simulator the way a flight simulator is one, built from what people learned flying the actual routes.

What "world model" means here

It means the architecture. One picture of your go-to-market that every plan starts from and every answer is read out of. What moves inside it is plain arithmetic you can read and change. It does not claim a trained model of your market.

How it checks itself

Three ways. Every run checks that cumulative wins and spend never go backwards. Pasting your own weekly wins fits it to your history and grades how close it gets. Then the loop back to reality, where you tell it what a play really did and its belief about that play moves.

The honest limits

  1. No effect size here is measured against a real install yet. The priors are practitioner heuristics written into the play library, every one printable and editable.
  2. The buyers are calibrated to B2B norms until you backtest on your own weekly wins and log outcomes.
  3. It compares arms on the same buyers. The revenue column is what 300 synthetic buyers did on illustrative numbers. It is never a forecast of yours.
  4. On this page no language model runs at all. Buyer answers are composed by the engine from the memory stream. Opened inside Claude, the same buyer answers in freer language. No language model ever decides a conversion, anywhere.
  5. Fifty plays run forward in the open simulator. The larger catalog inside the GTM OS demo is ranked and only part of it runs forward.
  6. What this browser learns stays in this browser. The loop across installs is the design-partner ask, which is roadmap.
  7. The state loading itself from your own sources is a GTM OS install feature, also roadmap. Here you set it with sliders.

Borrowed from

Park et al. 2023Wang et al. 2023GenBio AIDOMatrAIx

Run it yourself.

Everything here runs in your browser on an illustrative company. Nothing you type leaves it.

Buyer answers in free language, building a population from an ICP you describe, plus the written readout are the three things that need Claude. Everything else works without it.