Mission Intelligence Systems

Executive Organizational Diagnostic · Technical note

Scoring, interpretation and limits

This note describes what version 1.3 of the diagnostic computes, how it computes it, and what the resulting scores do not establish. It documents the production scoring engine rather than a separate conceptual model.

Technical-note version
1.3
Instrument fingerprint
8b9c5245b5a7 over 114 items
Effective date
July 30, 2026
Instrument state
Developmental; not empirically calibrated

Plain-language summary

The diagnostic is a self-report instrument with behaviorally anchored response options. It converts answers on a 1–5 ordinal scale into weighted construct means and composite metric scores. Those scores are directional indicators for reflection and discussion. They are not probabilities, causal estimates, clinical findings, audit evidence, or validated predictions of organizational outcomes.

Numerical operations on the response categories are a design convention used to summarize responses. They assume equal spacing between adjacent response values. That assumption has not been empirically established. A score of 80 is therefore not twice the organizational readiness of a score of 40, and a ten-point difference should not be interpreted as a measured ten-percentage-point change in performance.

Population, administration and question selection

The instrument is intended for executives and senior leaders describing current organizational conditions from their own perspective. It does not independently observe behavior or verify supporting evidence. Results may be affected by role, visibility, recall, response style, recent events and social desirability.

The production bank contains 114 questions across eight domains. A session selects 35–50 questions; the default is 40. Selection first applies any sector and organization-size eligibility rules, then greedily favors questions that improve coverage of constructs below the target of three observations. Remaining positions favor domain balance. Selected questions are interleaved across domains and their order varies by session.

Different respondents may therefore receive different item sets. Scores can be compared descriptively, but no equivalence study has yet established that every generated form has the same measurement properties.

Construct scoring

Every response is an integer from 1 to 5, with 5 always representing the healthiest described condition. Each item is tagged to two or more constructs with a weight from greater than 0 through 1. For construct c, the engine computes:

construct score(c) = Σ(item weight × response value) ÷ Σ(item weight)

The score is a rolling weighted mean on the 1–5 response scale. An item tagged to several constructs is an observation for each tagged construct. The current engine counts observations, but it does not estimate standard error, internal consistency, item discrimination or a confidence interval.

IDConstructClusterFour A engineFeeds metrics
A1Priority DisciplineAttentionattentionEAI, BI
A2Capacity ManagementAttentionattentionEAI, LL
A3Signal ClarityAttentionattentionEAI, SAC
A4Attention MarginAttentionattentionEAI, LL, EF
AL1Strategic CoherenceAlignmentalignmentSAC, BI, TR
AL2Cascade DepthAlignmentalignmentSAC, BI
AL3Incentive AlignmentAlignmentalignmentSAC, BI
AU1Decision ArchitectureAuthorityauthorityDV, AD, BI
AU2Authority CalibrationAuthorityauthorityDV, AD
AU3Governance EfficiencyAuthorityauthorityDV, EF, BI
AU4Escalation HealthAuthorityauthorityDV, AD, LL
AD1Psychological SafetyAdaptabilityadaptabilityRI, TR, ARP
AD2Learning LoopsAdaptabilityadaptabilityRI, TR, BI
AD4Change CapacityAdaptabilityadaptabilityTR, BI
F1Process LoadFrictionfrictionEF, LL, BI
F2Meeting EffectivenessFrictionfrictionEF, LL
F4Decision VelocityFrictionfrictionDV, EF, BI
RM1Risk Identification RigorRisk MethodologyattentionRI, BI
RM4Risk Ownership ClarityRisk MethodologyauthorityRI, AD
DG1AI ReadinessAI & DigitalalignmentARP, TR, BI

Metric construction and thresholds

A metric first takes the equal-weight mean of its observed input construct scores. Construct scores with no observations are excluded. Item weights affect the construct scores beneath a metric; the constructs themselves are not differentially weighted when combined into the metric.

This means a partially completed Builder Index is not merely a less complete estimate of a fixed composite. Its composition changes as previously unobserved constructs enter the mean. An early reading and a final reading can therefore differ because the set of included constructs changed, even before considering changes in the answers themselves.

Builder Index and AI Readiness Indicator values linearly rescale the construct mean from 1–5 to 0–100. The 0–100 format is an index without a percentage unit. Execution Friction and Leadership Load invert the healthy-oriented mean so that a lower displayed value is better. All other metrics remain on the 1–5 scale.

Bands and unlock thresholds are design thresholds. They have not been derived from outcome data, normative samples, sensitivity/specificity analysis or externally observed breakpoints.

Builder Index (BI)

Unlocks at question 15 and 15 observations

A 0-100 linear rescale of the equal-weight mean of observed input construct scores. Item weights apply within each construct; unobserved constructs are excluded until measured.

Input constructs
A1, A2, A3, A4, AL1, AL2, AL3, AU1, AU2, AU3, AU4, AD1, AD2, AD4, F1, F2, F4, RM1, RM4, DG1
Computation
Equal mean of observed input construct scores, linearly rescaled: ((mean − 1) ÷ 4) × 100.
Bands
Foundation: 0–39 · Developing: 40–59 · Operational: 60–74 · Advanced: 75–89 · High Performance: 90–100

Executive Attention Index (EAI)

Unlocks at question 5 and 3 observations

Degree to which leadership attention is concentrated, protected, and aligned with stated priorities.

Input constructs
A1, A2, A3, A4
Computation
Equal mean of observed input construct scores on the 1–5 response scale.
Bands
Critical: 1–1.9 · Fragmented: 2–2.9 · Developing: 3–3.9 · Disciplined: 4–5

Alignment Coherence (AC)

Unlocks at question 8 and 4 observations

Shared directional understanding across the leadership layer and into execution teams.

Input constructs
A3, AL1, AL2, AL3
Computation
Equal mean of observed input construct scores on the 1–5 response scale.
Bands
Fragmented: 1–1.9 · Surface: 2–2.9 · Emerging: 3–3.9 · Structural: 4–5

Authority Distribution (AD)

Unlocks at question 8 and 4 observations

Whether decision rights are distributed to the point of knowledge or concentrated in ways that create drag.

Input constructs
AU1, AU2, AU4, RM4
Computation
Equal mean of observed input construct scores on the 1–5 response scale.
Bands
Bottlenecked: 1–1.9 · Concentrated: 2–2.9 · Emerging: 3–3.9 · Distributed: 4–5

Decision Velocity (DV)

Unlocks at question 10 and 5 observations

Speed at which decisions reach resolution at the appropriate organizational level.

Input constructs
AU1, AU2, AU3, AU4, F4
Computation
Equal mean of observed input construct scores on the 1–5 response scale.
Bands
Stalled: 1–1.9 · Slow: 2–2.9 · Functional: 3–3.9 · High Velocity: 4–5

Execution Friction (EF)

Unlocks at question 12 and 5 observations

Total coordination overhead relative to execution value produced. Inverted - lower is better.

Input constructs
A4, AU3, F1, F2, F4
Computation
Equal mean of observed input construct scores, inverted for display: 6 − mean.
Bands
High Drag: 4–5 · Significant: 3–3.9 · Manageable: 2–2.9 · Low Friction: 1–1.9

Leadership Load (LL)

Unlocks at question 12 and 5 observations

Executive bandwidth consumed by reactive coordination and admin work vs. strategic leadership. Inverted - lower is better.

Input constructs
A2, A4, AU4, F1, F2
Computation
Equal mean of observed input construct scores, inverted for display: 6 − mean.
Bands
Overloaded: 4–5 · Strained: 3–3.9 · Managed: 2–2.9 · Protected: 1–1.9

AI Readiness Indicator (ARI)

Unlocks at question 20 and 6 observations

A composite index of the structural (not technical) conditions associated with AI investments producing organizational value. It is a rescaled mean of its input constructs, not an empirically calibrated probability of any outcome.

Input constructs
DG1, AD1, AU1, AD4
Computation
Equal mean of observed input construct scores, linearly rescaled: ((mean − 1) ÷ 4) × 100.
Bands
Low: 0–29 · Conditional: 30–59 · Favorable: 60–84 · Strong: 85–100

Risk Judgment (RJ)

Unlocks at question 20 and 6 observations

Risk methodology (identification, ownership, calibration) plus risk culture and leadership treatment of risk as judgment.

Input constructs
RM1, RM4, AD1, AD2
Computation
Equal mean of observed input construct scores on the 1–5 response scale.
Bands
Administrative: 1–1.9 · Emerging: 2–2.9 · Operational: 3–3.9 · Strategic: 4–5

Transformation Readiness (TR)

Unlocks at question 28 and 8 observations

A forward-looking indicator of readiness for large-scale organizational change, derived from your responses. It is not a validated prediction of whether a change effort will succeed.

Input constructs
AL1, AD1, AD2, AD4, DG1
Computation
Equal mean of observed input construct scores on the 1–5 response scale.
Bands
Unready: 1–1.9 · Fragile: 2–2.9 · Emerging: 3–3.9 · Ready: 4–5

Missing responses and evidence sufficiency

An unanswered item contributes nothing. Within a metric, an entirely unobserved construct is excluded rather than assigned zero. A metric becomes visible only after both its question-number threshold and minimum observation threshold are met. The adaptive session does not offer an “unable to observe” response or impute missing answers.

The self-directed short form does, since version 1.1. Its sixth response, “Insufficient evidence”, is stored as 0, is excluded from every condition mean, and is reported as evidence coverage: the share of the 24 items the respondent could substantiate. A condition scores only when at least three of its six items were substantiated; below that no index is computed and the reading names the conditions that lack evidence. Team Mode administration of the same 24 items is unchanged and offers no such response, so group aggregation, divergence and the widest-item selection are untouched.

Evidence sufficiency is answer coverage, not statistical confidence. The labels are Early for 0–9 answered questions, Developing for 10–20, Established for 21–34, and Complete for 35 or more. “Complete” means the coverage threshold was reached; it does not mean the result is definitive, accurate or validated.

Each metric card also has an evidence-fill display calculated as observations divided by twice the metric's minimum observation requirement, capped at 1. That display is likewise a coverage ratio, not a confidence interval or reliability estimate.

How the indicated constraint is selected

The self-diagnostic does not estimate a causal constraint. It ranks unlocked metrics for which a recommendation exists, excluding the Builder Index. Each metric is normalized to a 0–1 severity value: healthy-oriented metrics use (maximum − score) ÷ range; inverted metrics use (score − minimum) ÷ range. The metric with the highest resulting severity is displayed first and supplies the “Address First” recommendation.

All nine non-Builder-Index metrics currently have recommendation entries, so catalog coverage does not presently narrow the field. Unlock order does. Executive Attention Index can enter the ranking at question 5; Alignment Coherence and Authority Distribution at question 8; Decision Velocity at question 10; Execution Friction and Leadership Load at question 12; AI Readiness Indicator and Risk Judgment at question 20; and Transformation Readiness at question 28. Before a metric unlocks, it cannot be selected regardless of what later answers may show. The displayed result is therefore structurally biased toward earlier-unlocking metrics during partial completion.

Ranking also assumes that a linear normalization makes the different metric scales comparable. That is a design assumption, not an empirically established property. In particular, the operation treats distances between adjacent response-derived values as equal and treats equivalent normalized positions on different composites as commensurable. The current evidence does not establish either claim.

The phrase “primary constraint” should therefore be read as indicated area to investigate first, not as proof that the metric causes other organizational outcomes or that improving it will necessarily produce the greatest effect.

Team aggregation and divergence

Team reports require at least three respondents. For each unlocked metric, the engine reports the arithmetic mean, minimum, maximum, range and population standard deviation across respondents. Constructs receive the same descriptive summaries.

A divergence zone is flagged when the observed range exceeds 1.5 points on a 1–5 metric, or the equivalent 37.5 points on a 0–100 index. This is a design rule, not an empirically validated threshold. Divergence describes disagreement among respondents; it does not by itself establish that Alignment is the cause of the disagreement.

Intended and prohibited interpretations

Appropriate uses

  • Structure an executive conversation about current conditions.
  • Identify hypotheses and areas that warrant evidence gathering.
  • Surface disagreement among independently responding leaders.
  • Track directional change using consistent administration and context.
  • Prioritize questions for a facilitated diagnostic review.

Do not use it to

  • State the probability that an AI investment or transformation will succeed.
  • Claim that one condition caused an organizational outcome.
  • Rank employees, make employment decisions or evaluate individual performance.
  • Represent compliance, audit assurance or a professional risk opinion.
  • Benchmark organizations until an appropriate normative sample exists.
  • Treat small score differences as measured differences in capability.

Research use of responses

Since version 1.2 a participant in a team diagnostic may contribute their answers to research on the instrument. The offer is made after they have submitted and seen their own result, is refused at no cost, and is never made on the instruments that run entirely in the browser, whose pages promise that answers are never sent.

A contributed record holds the item responses, the instrument id and content fingerprint, sector, size and stage as bands, function and level from the fixed roster vocabulary, elapsed time, the month of contribution and an opaque organization key. It holds no name, no address, no title, no organization name, no team or response identifier and no timestamp for the answering itself. There is no path from a record back to a person. Retention is five years, and a withdrawal code shown once at contribution deletes the record on presentation; only a one-way hash of that code is stored.

This is what makes the analyses below possible and it does not perform them. Every analysis in the validation protocol needs item-level responses, and until this existed no store held any: team data is deleted at ninety days by promise, and the engagement archive deliberately keeps group means only. A corpus is a precondition, not a result, and nothing in this note's claims changes until an analysis has been run and reported here with its sample, its date and this fingerprint.

Validation status

Version 1.3 has automated tests for scoring behavior, question selection and aggregation. Those tests establish that the software follows its specified rules; they do not establish measurement validity.

No published study has yet established internal consistency, test–retest reliability, inter-form equivalence, factor structure, convergent or discriminant validity, criterion validity, predictive calibration, measurement invariance, normative bands or meaningful-change thresholds for this instrument. Until that work is completed, every output should be treated as developmental and directional.

The instrument is FROZEN for that work, and the fingerprint above is how a reader tells two versions apart. It is a content hash over every item's stem, construct tags and the text of all five behavioral anchors, so a reworded question or a redrafted anchor moves it even when the item count does not. A build in which the bank changes without this value and the version above changing with it fails, because responses collected either side of an unrecorded change cannot afterwards be pooled for any reliability or structure analysis, and nothing in the data would reveal that they had been. The analyses that will be run, the minimum sample each needs and, for each, the result that would count AGAINST the instrument are fixed in advance in the validation protocol held with this repository, along with the consent terms, the unit of analysis and the expert-review and cognitive-interview plans that come before any data at all. Fixing them before the first engagement is what stops an analysis chosen after seeing the data being reported as a test of it. Future technical-note versions will identify the data set, methods and results behind any validation claim. Changes to items, weights, constructs, thresholds or interpretation rules will be recorded in the version history below. What each analysis needs, and how much of that sample exists today, is published live at validation status, including the two thresholds that would fail: omega below 0.70 for a construct reported as a score, and an intraclass correlation below 0.60 for a measurement repeated with no intervention between.

Version history

1.2, 3 September 2026
An optional research contribution for team-diagnostic participants, described above, with its own consent, retention and withdrawal. Documentation and collection only: no item, weight, construct, threshold or interpretation rule changed, and the fingerprint above did not move.
1.1, 3 September 2026
The self-directed short form gains an “Insufficient evidence” response, excluded from every mean and reported as evidence coverage, with a condition scoring only when at least three of its six items are substantiated. The item bank did not change, so the fingerprint above did not move. Team Mode administration is unchanged.
1.0, 31 August 2026
The instrument frozen for validation: 114 items, fingerprint 8b9c5245b5a7, with the scoring, thresholds and interpretation rules this note describes.

Item-to-construct mapping

The table below is generated from the production question bank so that the published mapping cannot silently drift from the scoring engine. Weights apply within each construct's rolling mean.

ItemDomainQuestionConstruct weights
LEAD-01leadershipWhen a new high-priority demand arrives, what typically happens to the work already on the leadership team’s plate?A1: 1 · A2: 0.7
LEAD-02leadershipHow clearly do people below the executive team understand what leadership actually cares about most right now?A3: 1 · A4: 0.6 · AD1: 0.5
LEAD-03leadershipHow well do senior leaders protect time for the few things only they can do?A1: 0.8 · A4: 1
LEAD-04leadershipHow realistic is the leadership team about how much it can take on at once?A2: 1 · A4: 0.7
LEAD-05leadershipHow clear and consistent is the direction leaders give when priorities compete?A3: 1 · AL1: 0.6
LEAD-06leadershipHow safe is it for people to tell senior leaders something they don’t want to hear?AD1: 1 · A3: 0.5
LEAD-07leadershipHow disciplined are leaders about saying no to good-but-not-priority opportunities?A1: 1 · A3: 0.6
LEAD-08leadershipHow much of senior leaders’ calendars is consumed by meetings that don’t need them?A4: 1 · F2: 0.7
LEAD-09leadershipHow well do leaders develop the people beneath them rather than just direct them?AD1: 0.7 · AD2: 1
LEAD-10leadershipHow well do leaders delegate real authority rather than hold decisions themselves?A2: 0.8 · A1: 0.6
LEAD-11leadershipHow consistently do leaders model the behaviors they ask of everyone else?A3: 0.8 · AD1: 0.7
LEAD-12leadershipHow much of leadership’s time goes to proactive strategic work versus reacting to the day?A4: 1 · A2: 0.7
DEC-01decision makingFor recurring decisions, how clear is it who actually holds the authority to decide?AU1: 1 · AU2: 0.8
DEC-02decision makingWhen a decision needs to move up for resolution, what usually happens?AU4: 1 · AU3: 0.6 · F4: 0.7
DEC-03decision makingHow much does unclear decision ownership slow things down in practice?AU1: 1 · F4: 0.8
DEC-04decision makingHow well are people’s decision-making limits matched to what they can actually judge?AU2: 1 · AU4: 0.6
DEC-05decision makingHow much governance overhead sits on top of everyday decisions?AU3: 1 · F4: 0.7
DEC-06decision makingHow well designed are the forums and meetings where decisions actually get made?AU1: 0.8 · AU3: 0.8
DEC-07decision makingHow often are decisions made at the appropriate level rather than pushed up or down inappropriately?F4: 0.9 · AU2: 0.8
DEC-08decision makingHow often are made decisions reopened or second-guessed after the fact?AU4: 0.8 · AU1: 0.7
DEC-09decision makingHow well does the organization balance delegated authority with appropriate control?AU2: 0.9 · AU3: 0.7
DEC-10decision makingHow long does a typical significant decision take from raised to resolved?F4: 1 · AU4: 0.6
DEC-11decision makingHow well does the organization learn from decisions that turned out poorly?AU1: 0.7 · AD2: 0.9
DEC-12decision makingHow many approval loops sit between a routine request and getting it done?AU3: 0.9 · F1: 0.8
STR-01strategyIf you asked five senior leaders to state the strategy independently, how similar would the answers be?AL1: 1 · AL2: 0.8
STR-02strategyHow well do incentives and rewards line up with the stated strategic priorities?AL3: 1 · A1: 0.6
STR-03strategyHow clearly is the strategy expressed in terms people can act on?AL1: 0.9 · A3: 0.7
STR-04strategyHow well does the strategy translate into aligned goals as it cascades down the organization?AL2: 1 · AL3: 0.6
STR-05strategyHow disciplined is the organization about the trade-offs the strategy implies?A1: 0.8 · AL1: 0.9
STR-06strategyHow clear is the line of sight from frontline work to the strategy?AL2: 0.9 · A3: 0.7
STR-07strategyHow well does resource allocation actually follow the stated strategy?AL3: 0.9 · AL1: 0.7
STR-08strategyHow consistently is the strategy communicated over time?A3: 0.9 · AL2: 0.7
STR-09strategyHow well can the strategy adapt when conditions change without losing coherence?AL1: 0.9 · AD4: 0.7
STR-10strategyHow well does the organization distinguish true priorities from things that are merely important?A1: 1 · AL3: 0.6
EXE-01executionHow much of your teams’ capacity is consumed by process and meetings versus actual delivery?F1: 1 · F2: 0.9
EXE-02executionAfter an initiative finishes, what happens to what was learned?AD2: 1 · AD4: 0.7 · F4: 0.5
EXE-03executionHow often does process itself slow delivery of work that is already agreed?F1: 1 · F4: 0.7
EXE-04executionHow effective are the recurring meetings your teams rely on to coordinate?F2: 1 · A4: 0.6
EXE-05executionHow much capacity does the organization have to take on change while still running the business?AD4: 1 · AD2: 0.6
EXE-06executionHow smoothly does work move across handoffs between teams?F4: 0.9 · F1: 0.8
EXE-07executionHow well do teams turn what they notice during the work into improvements?AD2: 1 · F2: 0.6
EXE-08executionHow much accumulated process debt slows the organization down?F1: 1 · AD4: 0.6
EXE-09executionHow reliably do meetings end with clear ownership of what happens next?F2: 1 · F4: 0.6
EXE-10executionHow well does day-to-day execution stay connected to the strategy?AD4: 0.7 · AL1: 0.9
EXE-11executionHow does the organization respond when execution reveals that something isn’t working?AD2: 0.9 · RM1: 0.6
EXE-12executionHow much coordination effort does it take to get cross-functional work done?F1: 0.9 · F2: 0.8
RISK-01riskHow does the organization surface risks before they become problems?RM1: 1 · AD1: 0.6
RISK-02riskWho owns the organization’s most significant risks, and how clear is that ownership?RM4: 1 · AD2: 0.5
RISK-03riskHow rigorous is the organization at identifying the risks that actually matter?RM1: 1 · RM4: 0.6
RISK-04riskHow safe is it to raise a risk or bad news without being blamed?AD1: 1 · RM1: 0.6
RISK-05riskHow well is the authority to act on a risk matched to the person who owns it?RM4: 1 · AU4: 0.6
RISK-06riskHow well does the organization learn from near-misses, not just failures?RM1: 0.8 · AD2: 0.9
RISK-07riskHow does leadership treat risk - as a compliance chore or a judgment capability?AD1: 0.8 · AD2: 0.8
RISK-08riskHow well does the organization calibrate the likelihood and impact of risks?RM1: 0.9 · RM4: 0.7
RISK-09riskHow early does the organization tend to see trouble coming?RM1: 1 · AD1: 0.5
RISK-10riskHow consistently are significant risks reviewed and acted on over time?RM4: 1 · AD2: 0.6
RISK-11riskWhen a risk materializes, does the response focus on judgment or on blame?AD1: 1 · RM4: 0.6
RISK-12riskHow well is risk thinking built into major decisions rather than bolted on after?RM1: 0.9 · RM4: 0.8
TECH-01technology aiHow ready is the organization’s structure and culture to turn AI investment into real value?DG1: 1 · AD4: 0.7
TECH-02technology aiHow are decisions about technology and data actually governed?AU1: 0.9 · AD1: 0.6
TECH-03technology aiHow usable and trustworthy is the organization’s data for making decisions?DG1: 1 · AU1: 0.6
TECH-04technology aiHow prepared is the workforce for the ways technology will change their work?AD4: 1 · DG1: 0.6
TECH-05technology aiHow safe is it to experiment with new technology and sometimes fail?AD1: 1 · DG1: 0.6
TECH-06technology aiHow well does the organization learn and improve as it adopts new technology?DG1: 0.8 · AD2: 0.9
TECH-07technology aiHow well are technology decisions coordinated with the change management they require?AU1: 0.7 · AD4: 0.9
TECH-08technology aiHow healthy is the organization’s culture around using data and tools day to day?DG1: 1 · AD1: 0.5
TECH-09technology aiHow well does the organization scale a technology that works in one place to the whole?AD4: 0.9 · AU1: 0.7
TECH-10technology aiHow confident are you that current technology investments will deliver their intended value?DG1: 1 · AD4: 0.6
GOV-01governanceHow efficient is the organization’s oversight relative to the value it produces?AU3: 1 · F1: 0.8
GOV-02governanceHow well is governance designed as a performance enabler rather than a compliance exercise?AU1: 0.8 · A2: 0.6
GOV-03governanceHow clear are the roles and mandates of the organization’s governance bodies?AU3: 0.9 · AU1: 0.7
GOV-04governanceHow much reporting and administrative burden does governance place on the people doing the work?F1: 1 · A2: 0.6
GOV-05governanceHow well are controls sized to the actual level of risk?AU3: 1 · F1: 0.7
GOV-06governanceHow clearly do people understand how governance decisions get made and by whom?AU1: 0.8 · AU3: 0.8
GOV-07governanceHow much of leadership’s capacity is consumed by governance rituals versus real oversight?A2: 0.8 · F1: 0.8
GOV-08governanceHow well does governance speed up good decisions rather than slow everything down?AU3: 1 · AU1: 0.7
GOV-09governanceHow proportionate is compliance overhead to the risks it is meant to manage?F1: 1 · AU3: 0.7
GOV-10governanceHow well does executive and board oversight focus on what truly matters?A2: 0.8 · AU1: 0.8
CHG-01change transformationHow much capacity does the organization have to absorb and sustain major change?AD4: 1 · AL1: 0.6
CHG-02change transformationWhen change reaches the front line, how well does it actually take hold?AL2: 1 · AD1: 0.5 · AD2: 0.5
CHG-03change transformationHow well does the organization sustain change after the initial push fades?AD4: 1 · AD2: 0.7
CHG-04change transformationHow clearly is major change tied to a coherent strategic rationale?AL1: 1 · AD4: 0.6
CHG-05change transformationHow safe do people feel during periods of significant change?AD1: 1 · AD4: 0.6
CHG-06change transformationHow well does the organization learn and adjust as a transformation unfolds?AL2: 0.7 · AD2: 1
CHG-07change transformationHow well does the organization sequence multiple changes rather than pile them on at once?AD4: 1 · AL1: 0.6
CHG-08change transformationHow openly is resistance to change surfaced and worked through?AD1: 1 · AL2: 0.6
CHG-09change transformationWhat does the organization’s track record on major change look like?AD2: 0.8 · AD4: 1
CHG-10change transformationHow ready is the organization right now to take on its next major transformation?AL1: 0.8 · AD1: 0.7
LEAD-13leadershipWhat behavior do senior leaders visibly reward when they praise or promote someone?AL3: 1 · A3: 0.5
EXE-13executionWhat happens to someone who stops a project that is no longer worth finishing?AL3: 1 · A1: 0.5
GOV-11governanceHow closely do performance reviews and compensation follow the priorities leadership says matter most?AL3: 1 · AL1: 0.5
CHG-11change transformationWhen priorities change, how quickly do goals and incentives change with them?AL3: 1 · AL2: 0.5
LEAD-14leadershipWhen leaders hand over responsibility for something, how much authority goes with it?AU2: 1 · AU1: 0.4
EXE-14executionDo the people held accountable for a result hold the authority needed to produce it?AU2: 1 · AU1: 0.5
GOV-12governanceHow well do approval thresholds match the actual weight of what is being approved?AU2: 1 · AU3: 0.5
TECH-11technology aiWho decides where AI and new technology get used, relative to who understands the work?AU2: 1 · DG1: 0.5
DEC-13decision makingWhat happens to the work while a decision is sitting with someone above?AU4: 1 · F4: 0.5
RISK-13riskWhat does it cost someone to escalate a risk they cannot resolve themselves?AU4: 1 · AD1: 0.5
GOV-13governanceWhen something is escalated for a decision, what comes back?AU4: 1 · F4: 0.5
CHG-12change transformationDuring a major change, where do unresolved conflicts between teams go?AU4: 1 · AD4: 0.5
STR-11strategyHow many things does the organization currently call a top priority?A1: 1
LEAD-15leadershipHow much uncommitted time does a senior leader have in a typical week?A4: 1
EXE-15executionWhen two teams’ plans conflict, what do they resolve the conflict against?AL1: 1
CHG-13change transformationHow far down the organization can someone explain why the current change is happening?AL2: 1
GOV-14governanceFor the decisions this organization makes most often, is it written down who decides?AU1: 1
EXE-16executionOf the time a delivery team spends, how much goes to reporting and oversight rather than the work?AU3: 1
DEC-14decision makingIn a decision meeting, what happens when the most junior person disagrees with the most senior?AD1: 1
RISK-14riskAfter something goes wrong, what actually changes?AD2: 1
STR-12strategyWhen new work arrives and every team is already committed, what happens?A2: 1 · A1: 0.5
GOV-15governanceHow long does it take to get an answer to a question that is blocking work?F4: 1 · AU3: 0.5
GOV-16governanceWhen a new oversight or reporting requirement is introduced, is anything removed to make room for it?A2: 1 · AU3: 0.5
EXE-17executionWhen a team has to choose between two things leadership has asked for, do they know which one wins?A3: 1 · A1: 0.5
EXE-18executionBetween a piece of work being ready for a decision and the decision arriving, how much of that time is spent waiting?F4: 1 · AU4: 0.5
CHG-14change transformationDuring a change, how long does it take to reverse a decision that turns out to be wrong?F4: 1 · AD2: 0.5

Limits of use

The instruments are diagnostic tools. They are not guarantees of any outcome, and no result here should be treated as a plan that works because it was followed. Four limits govern every reading this platform produces, and they are stated in the same words on every result screen and in every report rather than summarized differently in each place:

  1. A reading is one vantage point, not the organization

    Every score here comes from what people could observe and were willing to report, from where they sit. A senior view stops where the information stops reaching it, and a team view is bounded by what the team has been told. That is a real signal about the organization, and it is not the same thing as ground truth. Treat a result as a claim to check against records, calendars, and the people closest to the work, not as a finding that has already been checked.

  2. A recommendation is a hypothesis, not an instruction

    What this produces is the most defensible next question given the pattern in the answers. It has no access to your funding cycle, your contracts, your regulator, your technology estate, or the person who is about to resign, and any one of those can be the real constraint while the instrument points somewhere else. The recommendation is worth acting on when your own judgment, and the evidence you can gather, agree with it. It is worth arguing with when they do not, and the disagreement is more useful than the score.

  3. Outcomes are decided in execution, by people and by systems

    Reading a diagnostic changes nothing. What changes an organization is a specific person with authority making a specific decision, and then the follow-through surviving contact with the work: competing priorities, staffing, incentives, contracts, data quality, system limits, vendors, and everything else outside these questions. Two organizations with identical results can end a year in opposite places, and the difference is what they did and what happened to them, not what they scored.

  4. Conditions move, so a result has a shelf life

    These readings describe a moment. A reorganization, a departure, a new system, or a change in demand can move a condition faster than any plan built on the old reading. Measure again rather than assuming a result still holds, and treat a number that has not moved as a question about whether anything actually changed.

The Builder Orientation Profile and the Builder Leadership Self-Check read a person rather than an organization, so their first limit is stated differently: what those instruments return is a disposition on the day it was answered, never a measure of ability, and never a basis for selecting, ranking or evaluating anyone. The other three apply unchanged.

These are held in lib/builder-intelligence/limits.ts at version 1.0 and are enforced by a build check, so a surface cannot ship a result without them and no two surfaces can state them differently.

Version history

Version 1.3 · July 30, 2026

Added the limits of use above as a single canonical statement rendered on every result surface and in all three reports. The limits were previously written per surface, differed between them, and were absent from the reports entirely, which is the artifact most likely to be read by somebody who never saw the instrument. Nothing about scoring changed and the instrument fingerprint is unchanged.

Version 1.2 · July 30, 2026

Initial public technical note. Documents the current question-selection, construct-scoring, metric-scoring, evidence-sufficiency, recommendation-ranking and team-divergence rules; states the instrument's developmental validation status and interpretation limits.

Read the practical methodologyOpen the diagnostic