Lead Scorer

Account Scoring in 2026: You Need 400 Scored Accounts Before Your Model Means Anything

Account scoring ranks companies by fit and intent instead of ranking individual leads. Here is the sample-size floor most teams miss, the four scoring approaches compared, and what to do before you have the data.

By Miljan @ Lead Scorer 10 min read

On 14 August, GTMnow published a detail from Salesforce that should unsettle anyone who runs a scoring model. Three quarters of Salesforce's inbound leads were never touched at all. Their 1-to-5 scoring worked exactly as designed, and the 1s and 2s became what their EVP of Agentforce Sales called "sawdust" — the pile nobody ever worked. Then they handed that pile to an engagement agent. It generated 100 million dollars in pipeline in 8 months.

The scoring model was not wrong about which accounts deserved a human. It was wrong that a low score meant no value. That distinction is the whole subject of this article.

The short answer

Account scoring ranks whole companies by how likely they are to buy, instead of ranking individual contacts. It rolls up fit signals (industry, headcount, official company data, tech stack) and engagement signals from every known person at that company into a single number, so you can decide which accounts get coverage rather than which inbox gets an email. It exists because B2B purchases are made by committees of 6 to 10 people, and a per-lead score structurally cannot see a committee forming.

The part almost nobody states: a scored tier only means something once you have roughly 200 accounts per tier with known outcomes. Below that, the difference between your "hot" and "warm" tiers is statistical noise you are staffing against.

Account scoring vs lead scoring, concretely

Lead scoring decides who gets contacted. Account scoring decides which companies get focused coverage. The failure mode of lead scoring in a committee purchase is specific and easy to picture: three mid-level engineers at the same company each read your pricing page twice this week. No single one of them clears a lead-score threshold. Rolled up to the account, that is one of the strongest signals you will get all quarter.

The inverse is just as common. One director opens four emails and clicks nothing, scoring high on engagement alone, while nobody else at that company has ever heard of you. Lead scoring promotes them. Account scoring does not. If you want the per-contact layer in depth, our guide to AI lead scoring covers it, and the MQL to SQL handoff covers what breaks between the two.

The four approaches, compared

"Account scoring" describes four quite different mechanisms that get sold under one name. They need wildly different amounts of data, and they fail in different ways.

ApproachMinimum data to workHow it failsBest fit
Explicit fit rulesNone — you write the criteriaEncodes your assumptions, including the wrong onesPre-product-market-fit and early outbound
Propensity model (ML)~400 scored accounts with outcomes, realistically 60+ closed-wonConfidently ranks noise when trained on thin historyTeams with years of clean CRM outcome data
Intent / third-party signalsVendor coverage of your categorySurges on competitor research and job seekersEstablished categories buyers actively search
LLM agent fit scoringA written ICP and verified company dataHallucinates firmographics if the data is not sourcedNiche ICPs with no historical volume

Most teams believe they are buying row 2 and are actually buying row 1 with a dashboard, or row 3 with a fit filter bolted on. That is not necessarily bad. It is only bad when you staff against it as though it were a validated prediction.

The 400-account floor (the calculation nobody publishes)

Here is the arithmetic that decides whether your model is real. Scoring splits accounts into tiers and claims the tiers convert at different rates. That is a two-proportion comparison, and it has a known sample-size requirement. At 95% confidence and 80% power, the accounts you need per tier is:

n ≈ 2 × (1.96 + 0.84)² × p̄(1 − p̄) ÷ (p₁ − p₂)²

Run it for the gap you are claiming:

Claim you are makingAccounts needed per tierTotal scored accountsClosed-won implied (at 15% win rate)
Tier A converts at 20%, Tier B at 10%~200~400~60
Tier A converts at 15%, Tier B at 12%~2,040~4,080~610

Read the second row again. Detecting a 3-point conversion difference — the kind of edge a scoring vendor will happily sell you — requires about 4,000 scored accounts with known outcomes. Almost no company under 10 million in revenue has that. Alex Vacca, who says he has audited 310+ B2B pipelines, puts the same conclusion in operator language: below 10 million, "you don't have enough closed-won deals yet for a scoring model to be right more often than it's wrong, so you buy speed instead of precision."

This is not an argument against scoring accounts. It is an argument against pretending a learned model is what you have. If you cannot fill the table above, use explicit fit criteria you wrote down and can audit — and be honest that they encode a hypothesis, not a finding.

What to score when you have no history

The workable substitute for a propensity model is a fit score you can defend line by line. Weight it toward fit, because intent data at low volume is mostly noise:

Score = (Fit × 0.6) + (Intent × 0.4)

Fit, rated 0-10, from inputs that are verifiable rather than inferred:

  • Industry classification and whether the company actually operates in your category
  • Headcount and revenue band, from filings rather than a scraped guess
  • Company age and legal status — a 4-month-old shell and a 20-year-old firm score differently
  • Whether the decision-maker role you sell to exists there at all
  • Technographic or operational evidence that your problem is present

Intent, rated 0-10, from what you can actually observe: aggregated engagement across all known contacts, hiring for roles that imply your problem, funding events, and public activity. Our ranked list of B2B buying signals covers which of these actually correlate with replies, and the intent data breakdown covers what third-party vendors can and cannot see.

The verification problem

An account score is a weighted sum of facts. If the facts are wrong, the score is worse than useless, because it is wrong with a confident number attached. This is the failure mode that has grown fastest as AI entered prospecting: a model asked to score a company will cheerfully invent a headcount, a funding round, and a CEO name, and the resulting score looks identical to a real one.

The fix is architectural, not a prompt. Every firmographic input should come from a source you can cite. This is the problem Lead Scorer's Outbound SDR agent is built around: it discovers companies through the web and the official French State registry (recherche-entreprises.api.gouv.fr, backed by SIRENE and INPI), so SIREN, legal status, company age and the registered dirigeant are read from filings rather than guessed. It then scores two levels — the company against your ICP, and the decision-maker within it — and rejects off-target accounts with a written reason you can argue with.

That written reason matters more than the number. A score of 7 tells a rep nothing. "Scored 7: matches ICP on sector and headcount, but the company was registered 5 months ago and has no operations history" tells them exactly whether to spend an hour on it. If you are comparing tooling in this space, Lead Scorer vs Clay covers the enrichment-first approach, and the lead scoring software comparison covers the rest of the category.

What to do with the low tier

Back to the sawdust. The structural error at Salesforce was not the scoring — it was treating "low score" as "no value" when what a low score actually means is "not worth expensive human minutes." Those are completely different instructions once the cost of a touch falls.

The rule that follows: score to allocate cost, not to allocate existence.

  • High tier — human research, multi-threading, custom sequences
  • Mid tier — agent-drafted personalised outreach, human approves before send
  • Low tier — automated nurture; revisit when a fit input changes, not on a calendar
  • Rejected — logged with the reason, so the rejection itself is auditable later

The fourth row is the one teams skip, and it is the one that makes scoring improvable. If you never record why an account was rejected, you can never find out that your rejection rule was wrong — which, given the sample-size table above, it probably is in at least one dimension.

How this fits the wider motion

Account scoring is a prioritisation layer, not a demand-generation strategy. It tells you where to spend the hours you already have. If the underlying target list is bad, a scoring model will faithfully rank a bad list. Start from a defined ICP, then score against it, then sequence. Teams running this as a full account-based motion should read the ABM playbook alongside this.

Lead Scorer runs the whole chain as one transparent pass — brief the agent in plain language, watch it discover companies from official data, review the two-level scores and their reasons, then approve the LinkedIn and email drafts (each one reviewed by a second model before it reaches you). Plans start at 49 euros a month on the pricing page.

The honest summary

Account scoring is the right unit of analysis for B2B — committees buy, individuals do not. But the industry sells it as a predictive model to companies that mathematically cannot have one yet. If you have 400 scored accounts with outcomes, build the model. If you do not, write explicit fit criteria, source every input from something verifiable, keep the reasons attached to the scores, and never confuse a low score with a worthless account. Salesforce found 100 million dollars in the pile they had already written off.

Further reading: Lead qualification frameworks · B2B buying signals ranked by strength · The AI lead scoring guide.

Frequently asked questions

What is account scoring?

Account scoring ranks whole companies by how likely they are to become customers, instead of ranking individual contacts. It aggregates fit signals (industry, size, tech stack, official company data) and engagement signals across everyone at that company into one number, so sales can decide which accounts get coverage rather than which inbox gets an email.

How is account scoring different from lead scoring?

Lead scoring decides who gets contacted. Account scoring decides which companies get focused coverage and coordinated plays. In a committee purchase with 6 to 10 people involved, a single lead score is misleading: three mid-level people browsing pricing is a stronger buying signal than one director opening an email, but lead scoring cannot see that pattern because it never rolls contacts up to the company.

How many accounts do I need before an account scoring model works?

Roughly 200 accounts per tier with known outcomes, so about 400 in total, and that is only enough to prove a large gap (a tier that converts at 20% versus one at 10%). To prove a subtler 3-point gap you need about 2,000 per tier. Below that, the tier differences you see are noise, not signal.

What is a simple account scoring formula?

A workable starting formula is Score = (Fit x 0.6) + (Intent x 0.4), where Fit is a 0-10 rating against explicit ICP criteria you wrote down, and Intent is a 0-10 rating of recent observable activity. Weight fit higher than intent early on: intent data is noisy at low volume, and fit criteria are auditable while a propensity model is not.

Should low-scoring accounts be ignored?

No, and this is the most expensive mistake in the discipline. Salesforce found that its 1s and 2s, the pile nobody ever worked, produced 100 million dollars in pipeline in 8 months once an agent picked them up. A low score should mean low priority for expensive human time, not deletion.

Can AI do account scoring without historical data?

Yes, but not the same way. A machine-learning propensity model needs closed-won history. An LLM-based agent instead evaluates each company against an ICP you describe in plain language, using verified company data, and returns a score with a written reason. It is a fit judgment, not a statistical prediction, which is exactly what you want when you do not yet have 400 scored accounts.

What data should feed an account score?

Firmographics (industry code, headcount, revenue band, location), official registry data such as legal status and company age, technographics, hiring signals, funding events, and aggregated engagement from every known contact at the account. The critical rule is that each input must be verifiable, because an account score built on hallucinated firmographics is worse than no score at all.

How often should account scores be recalculated?

Fit scores change slowly and can be refreshed monthly or when registry data changes. Intent scores decay fast and should be recalculated weekly at minimum, ideally on each new signal. Mixing a stale intent score into a live priority list is the most common reason reps stop trusting the number.

Keep reading