A GTM simulator that makes judgment calls.
A world model of your go-to-market. Aerospace simulates before it flies. Drug discovery simulates cells before it touches a patient. Go-to-market still tests in production, on your budget, your reputation and your buyers' patience. The simulator is where you run the experiment first. It tells you what to run, what to kill and when you will know.
It tells you what is going to work next.
Ten minutes of setup, then you can try any plan you like before you spend a euro on it.
You describe your company
What you spend in a week. How many deals your team can work at once. How many of your buyers have heard of you. What you already run today. Eight numbers and a list.
It plays the next four months
Three hundred buyers, one week at a time, twenty-four times over so you see a range instead of a single guess. Then it does the same for every plan you want to try, on the same buyers with the same luck, so the plan is the only thing that changed.
You get a ranking and a reason
Which plan came out ahead most often, by how much, the week you would be able to tell in real life, plus what stopped the ones that failed. It ranks your options. Treat the euro amounts as illustrative until you feed it your own history.
Four questions every GTM leader asks. Three more they should.
What is going to work?
Your plans, ranked by how often each one finished ahead of doing nothing. Ahead in 70 runs out of 100 or better and it says works. Between 55 and 70 it says coin flip, because that is what it is.
What is not going to work?
The same list read from the bottom, at 55 out of 100 or worse, with the thing that killed each plan named out loud. You ran out of salespeople. You ran out of budget. Too many people had to say yes. The price was wrong. The buyers got tired of hearing from you.
What is my next best step?
The one play worth doing now, because it gains you the most or because it teaches you the most. You pick what to try and it shows you where that lands against everything else the same hours could have bought. Sometimes the honest answer is to stop doing something you already do.
What do I do after that?
Name the number you want to hit and the most you will spend in a week. It works backwards to find the order of moves that gets you there, or it tells you the target is out of reach on that money. Most teams start here, because they wake up wanting a number to move.
The three that rarely get asked
When will I know?
The first week the early signs pull away from doing nothing and stay there. Write that week down before you start, then stop if it comes and goes.
What would I have to believe?
Which single assumption, if it turns out to be 30 percent off, changes the whole ranking. That is the number worth going and measuring first.
How much should I trust this?
Paste your real wins week by week and it grades how close it gets to your own history. Tell it what a play actually did and it adjusts what it believes about that play. Until you do one of those, trust the ranking and the reasons behind it. The euro amounts are illustrative.
Same buyers. Same dice.
The simulator calls each copy of your company an arm, the way a drug trial does. Everything below is two arms.
Three hundred buyers, drawn once. Both copies of the company meet the same three hundred, in the same order, with the same luck. Whatever separates them at the end came from the plays.
buyers won in A only
The gap is the change. Nothing else was different.
Move a slider. Watch the answer change.
It has already worked out what is holding this company back, picked the three plans most likely to fix it and run them. Change anything and it runs again.
Five steps, one loop.
State
Weekly budget, how many deals your team can work at once, outbound and marketing hours, how well known you are, how much you are trusted, how tight buyer budgets are, how hard competitors push. Those sliders are your company as it stands.
The plans being compared
Make copies and change one thing in each. Copy A goes hard on outbound. Copy B fixes conversion. Both start from the same day, so any difference at the end came from the change.
Play them forward
Each copy runs week by week. Reps fill up. Lists run down. Deals take their real time to close. Plays that do not fit the weekly budget pause that week.
The readout
Every plan is scored against doing nothing, because a plan only counts if it beats that. You get a direction, a range and a plain answer.
Log reality
Tell it what really happened when you ran a play. It adjusts what it believes about that play and the next run reorders. What this browser learns stays in this browser.
In a drug trial one group gets the drug, one gets a placebo, everything else is kept identical. Same idea here, with your company in place of patients. The pill is a budget decision.
One living picture, built from four layers.
A picture of your go-to-market that plans can be tested against.
Your state
Eight sliders. Weekly budget, deals worked at once, outbound hours, marketing hours, brand awareness, baseline trust, buyer budgets, competitor pressure. Plus the plays you already run.
The buyers
Three hundred of them, each with a memory of every touch you send and a weight for how much it mattered. A peer referral lands heavier than an ad. Enough noise and they get tired, then they opt out.
The play library
Fifty plays, each with a weekly cost, what it eats from your shared hours, how long before anything shows, plus a starting assumption about how much it helps. Every one runs forward in time.
The learning loop
Tell it what a play really did and it adjusts what it believes about that play, so the next run reorders on your evidence. Paste your own weekly wins to see how far it sits from your history.
Ask a simulated buyer why.
Pick one of the buyers you lost and ask. The answer is composed from that buyer's own memory, so it names the touches, the weeks and the reason.
Why time changes the answer.
A list of good ideas puts them in order. It cannot tell you the third one breaks the second, or that the first has no chance of showing up inside the quarter you judge it in.
Seller capacity
A team can only hold so many first meetings a week. Anything booked past that ceiling sits waiting and goes cold.
List saturation
The further you work through your list, the worse each email does. Sending twice as much buys nowhere near twice the replies.
Pipeline lag
Most of what this quarter builds closes next quarter.
Compounding content
Pages take months to warm up. The same play reads flat over 90 days and strong over a year. That is the calendar talking. The play is fine.
Shared pools
Hours, budget and the list are all finite. Every play eats from the same pools, so two email plays on one list quietly starve each other.
Tired buyers
Keep touching the same people and they stop noticing, then they opt out. A plan that looks strong on volume can lose on how people feel about it.
Payback against the clock
Every plan prints what a customer cost you and how many months of margin it takes to pay that back, so a move that works and pays back too late is visible as that.
Median mid-market deal in the run above. Most of what this quarter builds closes next quarter.
What is in there and what is not.
The simulator is the engine. The GTM OS is the harness around it.
What it is built from
The rules inside are written-down practitioner experience, the play library above. It is a simulator the way a flight simulator is one, built from what people learned flying the actual routes.
What "world model" means here
It means the architecture. One picture of your go-to-market that every plan starts from and every answer is read out of. What moves inside it is plain arithmetic you can read and change. It does not claim a trained model of your market.
How it checks itself
Three ways. Every run checks that cumulative wins and spend never go backwards. Pasting your own weekly wins fits it to your history and grades how close it gets. Then the loop back to reality, where you tell it what a play really did and its belief about that play moves.
The honest limits
- No effect size here is measured against a real install yet. The priors are practitioner heuristics written into the play library, every one printable and editable.
- The buyers are calibrated to B2B norms until you backtest on your own weekly wins and log outcomes.
- It compares arms on the same buyers. The revenue column is what 300 synthetic buyers did on illustrative numbers. It is never a forecast of yours.
- On this page no language model runs at all. Buyer answers are composed by the engine from the memory stream. Opened inside Claude, the same buyer answers in freer language. No language model ever decides a conversion, anywhere.
- Fifty plays run forward in the open simulator. The larger catalog inside the GTM OS demo is ranked and only part of it runs forward.
- What this browser learns stays in this browser. The loop across installs is the design-partner ask, which is roadmap.
- The state loading itself from your own sources is a GTM OS install feature, also roadmap. Here you set it with sliders.
Borrowed from
Every term on this page, in plain words.
One definition, used on the landing page, the explainer and the runner. Written for a marketer who has never opened a terminal.
- World model
- A running picture of your company that the simulator keeps updating week by week. Everything you try starts from that same picture, so two plans can be compared fairly.
- State
- Everything true about your go-to-market right now. How many sellers you have, how big your list is, what a deal is worth, how long one takes to close, what is in the bank. The starting point for any test.
- Arm
- One version of your plan. If you are deciding between spending the quarter on outbound or on fixing conversion, that is two arms. The word comes from clinical trials, where each group of patients is an arm of the study.
- Branch
- Making a copy of your company on the same day, then changing one thing in the copy. Both versions run from the identical starting point, so any difference at the end came from the change you made.
- Play
- One concrete thing you could do. Score accounts by intent. Publish comparison pages. Hire a rep. There are 54 in the library, and the simulator can run every one of them forward in time.
- Run
- One pass through the next few months, a week at a time. The simulator does 24 of these by default, and you can dial that from 4 up to 96, changing the guesses slightly each time, so you see a range of outcomes instead of one lucky answer.
- Readout
- What comes back after a run. Pipeline, meetings, revenue, cash. Every readout is measured against what would have happened if you did nothing.
- Doing nothing
- The same company over the same months with no new plays at all. Every arm is scored against this, because a plan only counts if it beats leaving things alone.
- Assumption
- A number the simulator needs that you have not measured. How many meetings one seller can hold, how fast a page ramps. Each one is shown with a range. You can type over it.
- Where a number came from
- Green means the number is yours, read from your systems or typed in by you. Amber means it is a sensible guess with a range around it. Red means nobody knows, so anything resting on it cannot be called.
- Default
- A sensible guess used when you have not measured something yourself. Always shown in amber with the range it could sit in, plus a note on how you would find the real number.
- Ceiling
- A limit that stops a plan from growing any further. Two reps can only hold so many meetings a week. A list of 1,800 accounts runs out. Hitting a ceiling says the plan ran out of room. The next move is to raise it.
- List saturation
- The more of your list you have already contacted, the worse the next email does. Sending three times as much does not get three times the replies, because you are working further down the same finite list.
- Pipeline lag
- The gap between creating an opportunity and closing it. If your deals take 75 days, most of what you build in a quarter closes after that quarter has ended. Judge the quarter on revenue alone and every plan looks like a failure.
- Agreement
- Out of 100 runs with slightly different guesses, how many times this plan beat doing nothing. High agreement means the answer holds up even when the guesses move. It is not a probability that the market will comply.
- Sensitivity
- Which unknown number is doing the most work in the answer. If an amber guess sits at the top of this list, that is the thing worth measuring before anyone spends money.
- Shared pool
- Something finite that every play draws from. Hours in the week, budget, the list itself, your email reputation. Two plays can quietly starve each other by pulling on the same pool.
- Horizon
- How far ahead you are looking. The same play can read flat over 90 days and strong over a year, so the window you pick changes the answer as much as the plan does.
- Forward mode
- You pick a move, and the simulator tells you where it lands against everything else you could do with the same hours.
- Backward mode
- You name the number you want to move. The simulator works back to the plays that act on it. Most teams start here, because they wake up wanting a number to change.
- Verdict
- The one-line answer for each arm. Works, likely, works until a certain week, cannot tell yet, or will not work. The rule underneath matters. A plan is only failed when the thing that fails it is a number you gave.
- Payback
- How long a move takes to earn back what it cost. A plan that works but pays back after the money runs out did not work.
- Deliverability
- Whether your emails actually land in the inbox. Send too often to the same list and this falls, which quietly drags down every email play you run.
- Ramp
- The time between starting something and it working at full strength. A new rep takes months to reach full output. A new page takes months to reach full traffic.
- Tier-1 list
- The accounts that actually fit what you sell. Once you have worked through them, the next batch converts worse, which is why list size sets a hard limit on outbound.
- Prerequisite
- Something a play cannot run without. A call-coaching play needs call recordings. If the source is missing, the simulator says so instead of pretending the play would work.
- Why more than one run
- Any single run depends on guesses that could be off. Running it 24 times over with the guesses nudged each time shows the spread of what could happen. The width of that spread is the honest part.
- Signal
- Anything that hints a company is in the market right now. A hiring post, a funding round, someone visiting your pricing page twice, a competitor mention. A signal engine watches for these and tells you who to call today.
- Scoring engine
- Ranks your accounts and leads so reps work the best ones first. It reads who fits what you sell and who is showing interest, then puts a number on each. Without it, reps work whatever is at the top of their list.
- Routing
- Deciding which rep gets which lead. Speed counts as much as the match here. Bad routing is one of the quietest killers in go-to-market, because a good lead sitting unassigned for three days converts far worse than the same lead worked in an hour.
- Lead engine
- The machinery that turns strangers into named people you can contact. Finding them, checking they fit, enriching what you know about them, then handing them to the right place.
- Outreach engine
- The machinery that contacts people and handles what comes back. Sequences, replies, follow-ups, booking the meeting. It is the part most teams mean when they say outbound.
- ABM
- Account-based marketing. Instead of casting wide and seeing who bites, you pick a specific list of companies you want and aim everything at them. Common when deals are large and the buyer list is short.
- Lifecycle
- Everything that happens after someone becomes a customer. Onboarding, staying active, renewing, expanding, or churning. Usually where the cheapest growth is hiding.
- GEO
- Getting mentioned by AI assistants when someone asks them what to use. Like SEO, except the thing you are trying to rank in is ChatGPT or Claude instead of a page of blue links.
- Enrichment
- Filling in the blanks on a record. You have an email, and enrichment adds the company size, the industry, the job title, the tech they run, so scoring and routing have something to work with.
- CAC
- What it costs you to win one customer, everything included. If CAC goes up while deal size stays flat, the machine is getting less efficient even when revenue looks fine.
- ICP
- Ideal customer profile. The kind of company you actually win, keep and make money on. Written down properly it decides who you contact, how you score them and what good looks like.
- Pipeline
- The deals currently in progress, plus what they are worth. It is the earliest honest read on whether a plan is working, because revenue arrives too late to steer by.
- Exponent
- A dial for how sharply something falls away. At 1 the drop is steady. Above 1 the first contacts are worth much more than the last ones. Below 1 the list holds up better as you work through it.
- Lognormal sigma
- How spread out your deal closing times are. A small number means most deals close near the median. A large number means some close fast and a long tail drags on for months.
- Share lost per week
- The slice of untouched interest that goes cold every week nobody works it. At 0.45, a booked meeting nobody chases loses roughly half its value each week it sits.
- Engine and harness
- The simulator does the thinking. The GTM OS around it holds your data, keeps the picture current and remembers what you already tried, so you never retype anything.
Run it yourself.
Everything here runs in your browser on an illustrative company. Nothing you type leaves it.
Buyer answers in free language, building a population from an ICP you describe, plus the written readout are the three things that need Claude. Everything else works without it.