Reading evidence weight: telling a proven capability from a claimed one
Every capability statement a firm writes says it can do everything. Weighting each claim by the evidence behind it is what makes a capability record able to discriminate at all.
Evidence weight is how much proven delivery stands behind a stated capability. A capability supported by four completed engagements carries more weight than one supported by a single mention, and one with no delivery behind it carries none. Ranking by evidence lets a capability record discriminate without anyone self-assessing.
Open any firm's capability matrix. Services down one axis, sectors across the other, and hardly a blank cell in it.
A matrix with no zeroes carries no information. If the firm can do everything for everyone, the document has told you nothing you did not know before opening it, and worse, it cannot be used for the one thing a capability record is for: choosing between two options.
Why claimed capability always inflates
This is not dishonesty. It is the predictable output of how capability gets recorded.
Self assessment rewards breadth. In a firm where staffing follows stated capability, describing yourself narrowly is career limiting. Everyone therefore reports the widest defensible range, and they are all being truthful: the tax partner really has touched transfer pricing. The result is a directory where forty people match every search, which discriminates exactly as poorly as a directory where nobody does.
Marketing material is written to win, not to describe. A capability statement is a sales document. It is supposed to present the firm at its strongest, and it does. It is the wrong artefact to run internal resourcing from, and it is frequently the only one that exists.
Nothing is ever removed. Capability accumulates and never decays. A practice the firm exited in 2019 is still on the list, because there is no process that takes things off.
A capability list that cannot say no about anything cannot be relied on about anything.
What the alternative measures
The alternative is not to ask better questions. It is to stop asking, and infer capability from delivery.
Every engagement the firm has completed is evidence that the firm can do the thing that engagement was. If you record engagements properly, with the people who delivered them and the capability each one demonstrates, capability stops being a claim and becomes a consequence. Nobody rates themselves. Nobody is ranked against a colleague by a human being. The record simply reflects what has actually been done.
The expert-finding literature has converged on this over two decades. The systematic reviews land in the same place repeatedly: expertise should be inferred from evidence of work, and recommendations should be supported by the documents used to make them.
What weight is, concretely
Once capability is inferred from delivery, the number of pieces of evidence behind each claim becomes meaningful, and it can order things.
In OrgAtlas each item carries a gauge showing how much proven evidence stands behind it, and that gauge sets the running order within its group. The effect in use is that the person who has genuinely led four engagements in a domain sits at the top without anyone deciding they should, and the person copied on one sits where the evidence puts them.
Three details matter more than the mechanism.
Role is part of the weight. Leading an engagement is different evidence from being copied on it. Frequency of appearance is a poor proxy for contribution, and any weighting that ignores role reproduces the failure of counting emails.
Distinct clients count for more than repetition. Four engagements for one client evidences a deep relationship. Four across four clients evidences a transferable capability. They are different claims and should not weigh identically.
Weight is not a score out of ten. It is a reading of how much stands behind something, and it is always openable. The number is a way in to the evidence, not a replacement for it.
Drawing the hollow ones
The most useful part of this is the part that looks like an omission.
A capability with no delivery behind it is drawn hollow rather than left out. That is a deliberate choice with a real cost: it means the atlas visibly displays things the firm cannot prove it does, which is uncomfortable in a way that a clean list is not.
It is worth the discomfort for two reasons.
For resourcing, a hollow capability is a warning. Somebody about to staff an engagement can see that the firm's claim rests on nothing, before rather than after committing.
For business development, a hollow capability against a client who demonstrably needs it is the definition of whitespace. If capabilities with no evidence were simply absent from the record, that gap would be invisible, and the firm would keep discovering it the expensive way, when the client buys it from somebody else.
Absence is a finding. A record that only shows what you have done cannot show what you have not.
The credibility argument
There is a deeper reason to show weight rather than assert capability, and it comes back to the theory the whole product is built on.
Credibility, in transactive memory, is whether members of a group trust one another's knowledge enough to rely on it rather than redo it. In a large firm, trust travels along the lines the work travelled, so it is dense within a practice group and thin between them. A partner will rely on a colleague they delivered with in 2019 and independently rebuild the work of a colleague two floors up they have never met.
Evidence weight is how that trust gets extended to strangers. A partner does not need a prior relationship with the person the atlas surfaces if they can see the four engagements behind the recommendation and open any of them. Credibility moves from being a property of relationships, which does not scale, to a property of the record, which does.
That is what the gauge is for. Not to score people. To let a colleague who has never met them decide whether to rely on them.
Sources
This article stands behind sheet 03 on the homepage, The live demo.
Knowledge graphs, explained
Why the questions firms ask about their own work are relationship questions, why relationship questions need a graph, and what that means in practice rather than in architecture diagrams.
D2GraphRAG versus vector search
Vector retrieval is very good at finding passages that resemble a question. Relationship questions do not resemble their answers, which is why firms get fluent nonsense back.
D4Mapping an account, worked
What actually happens between forwarding a year of correspondence and having an account you can navigate. Walked through on the demo account, step by step, including what it cannot do.
See the argument running.
The homepage carries a working atlas of a demo account, with every fact tied back to the document it came from. Or book a call and we will walk through an account you know.