A Companion Paper to the Production Readiness Assessment
How I decide what can hold.
The methodology behind the verdict — the dimensions I score, the rubrics I use, the gates that force a Stop, and the logic that turns scoring into a defensible recommendation.
Argyris Skouloudis
Senior operator-advisor · Enterprise AI
Athens · 2026
argyris@skouloudis.com
§ 01
The operating principles
Before the scoring, six commitments.
A scoring framework is only as honest as the principles behind it. These six govern how every assessment is run.
i.
Score before deciding
Every initiative is scored across the six dimensions before any verdict is reached. Pattern recognition follows analysis, not the other way around.
ii.
Anchored rubrics, not gut feel
Scores use written behavioural anchors. Two senior reviewers scoring the same initiative should land within one point on each dimension.
iii.
Stop must be reachable
Hard gates on three dimensions allow Stop and Convert-to-Automation verdicts on principled grounds. The math has to make these verdicts possible.
iv.
Every score has a rationale
No score travels alone. Each receives a written one-paragraph rationale referencing the evidence that drove it. The trail is the deliverable.
v.
Judgment overlays math
Scoring produces a candidate verdict. Final verdicts incorporate organisational context, time-to-value, vendor execution risk, and political reality. The framework is a discipline, not an oracle.
vi.
Independence in writing
Where a recommendation touches platforms I am commercially involved with, the conflict is disclosed in writing and the client may request an independent second opinion.
§ 02
The six dimensions
What I score, in order.
Six dimensions, scored independently, weighted equally in the headline total — but three of them carry gate authority that can override the total. The dimensions are sequenced so that the earlier ones fail faster: if Workflow Suitability or Knowledge Readiness is broken, no amount of governance design or KPI thinking will save the initiative.
Most enterprise AI initiatives do not fail because the model is weak. They fail because no one designed for operational trust before deployment — and the failure was visible in the scoring before the first line of code.
— The operating thesis
Is the work itself a fit for AI execution — or would deterministic automation be more reliable, cheaper and faster to audit?
Sub-criteria
Ambiguity present in inputs·Exception frequency·Contextual reasoning required·Variability across cases·Need for human judgement
0
Wrong fit
Workflow is deterministic, rule-based, repeatable. Zero ambiguity. AI adds cost and risk without producing value that automation could not.
1
Likely wrong fit
Mostly deterministic with rare edge cases. Conventional automation with rule-based exception handling is more reliable and cheaper to operate.
2
Mixed
Some genuine ambiguity, but routine cases dominate. Hybrid approach viable — automation for the bulk, AI for the long tail.
3
Strong fit
Substantial ambiguity, contextual reasoning required across most cases. Variability is too high for rule-based automation to cover.
4
Native fit
High ambiguity, heavy context-dependence, novel situations common. The work is what AI is genuinely better at than any alternative approach.
Gate behaviour
A score of 0 or 1 forces the verdict toward Convert to Deterministic Automation or Stop, regardless of strong scores elsewhere. You cannot agentify your way out of a workflow that should never have been agentified.
Is the underlying information production-grade — or is it documents pretending to be knowledge?
Sub-criteria
Source authoritativeness·Semantic consistency·Structural quality·Retrieval testability·Tribal-knowledge dependency
0
In people's heads
Knowledge lives with individuals. No retrievable corpus exists. Capturing it is itself a multi-quarter project before any AI work can begin.
1
Fragmented
Conflicting truth sources, inconsistent terminology, undocumented exceptions. Documents exist but cannot be relied upon. Heavy remediation required.
2
Uneven
Substantial documentation exists; quality varies by domain. Gaps and redundancies present. Selective curation needed before deployment.
3
Production-capable
Well-structured corpus with consistent terminology and identifiable authoritative sources. Minor gaps acknowledged and tracked.
4
Production-grade
Versioned, governed, retrievable, semantically consistent. Gaps known and documented. Knowledge operations function as a discipline, not as a side-effect of other work.
Gate behaviour
A score of 0 or 1 forces the verdict toward Redesign or Stop. You cannot solve a knowledge problem with a model — the model amplifies whatever it retrieves, including the noise.
Can this hold under audit, compliance review, and the human-oversight requirements implied by the workflow's risk profile?
Sub-criteria
Auditability of decisions·Provenance & traceability·Access-control fit·Policy enforceability·Regulatory exposure clarity
0
Absent
No governance framework exists or can be defined within the workflow's risk profile. Exposure is structural, not addressable.
1
Insufficient
Severe gaps. Regulatory or compliance exposure exceeds organisational tolerance. Existing controls do not extend to AI behaviour.
2
Achievable
Governance achievable but requires substantial design work and cross-functional alignment. Risk, compliance, and IT have not yet aligned on requirements.
3
Clear
Governance requirements clear, achievable, and largely aligned with existing controls. Owners identified for each control area.
4
Audit-ready
Governance fully mapped. Controls validated against EU AI Act, GDPR, DORA, sector regulators. Audit-trail design ready before deployment.
What breaks when confidence drops — and is the fallback designed, owned, and tested before deployment?
Sub-criteria
Exception detection logic·Escalation paths·Confidence thresholds·Human-in-the-loop triggers·Graceful degradation behaviour
0
Not considered
No failure-path thinking has occurred. The system is assumed to always work. Production exposure is open-ended.
1
Acknowledged only
Failure paths acknowledged but not designed. Escalation is "send to a human" without specifying which human, on what trigger, in what time window.
2
Partial
Basic failure paths defined; gaps remain in confidence thresholds, escalation routing, or graceful degradation behaviour.
3
Designed
Failure paths designed, escalation owners identified, confidence thresholds calibrated to workflow risk profile.
4
Designed & tested
Comprehensive: confidence-driven routing, defined intervention triggers, tested fallback behaviours, shadow-mode validation complete.
Can this actually operate inside the real systems-of-record — or does it sit beside them as a glorified search box?
Sub-criteria
Systems-of-record integration·API & data-flow viability·Ownership of operational running·Downstream-process impact·IT readiness
0
Blocked
Required systems are inaccessible, undocumented, or institutionally blocked. Deployment is structurally impossible in the current state.
1
Fragile
Integration technically possible but operationally fragile. No clear ownership of running the system once deployed.
2
Achievable
Integration viable with moderate effort. Ownership boundaries between IT, operations, and the AI function need explicit negotiation.
3
Ready
Systems accessible, integration patterns identified, ownership of operational running clearly assigned.
4
Operationally complete
Integration well-defined, operational ownership assigned, downstream processes mapped, change-management implications understood.
Gate behaviour
A score of 0 forces the verdict to Stop, regardless of other dimensions. A score of 1 forces the verdict toward Stop or Keep Human-Assisted. AI without operational integration is expensive search.
Will value actually be measurable — or are we agreeing to deploy on faith and review on vibes?
Sub-criteria
Baseline metrics exist·Target deltas quantifiable·Attribution model viable·Adoption signals trackable·Value timeline realistic
0
Asserted
No baseline, no metric, no attribution model. Value is asserted in the business case rather than measurable in the operational data.
1
Vague
Directional metric ("improve customer experience") with no baseline and no attribution. Result will be unfalsifiable in either direction.
2
Partial
Metric defined but baseline measurement gaps remain. Attribution model requires design before deployment.
3
Measurable
Baselines exist, targets defined, attribution model viable. Adoption tracking instrumented.
4
Provable
Full measurement design: baselines captured, targets quantified, attribution model proven elsewhere, value timeline realistic against operational change rate.
§ 03
The scoring system
Five anchors, not five gut feels.
Each dimension is scored on a 0–4 scale using behavioural anchors — written descriptions of what each score actually looks like in the real organisation. The anchors exist so that scoring is reproducible: two senior reviewers assessing the same initiative should land within one point on each dimension. If they cannot, the rubric is wrong, not the reviewers.
Total possible score across the six dimensions: 24 points. The headline number matters less than the distribution — a 16 made up of strong governance and weak workflow fit produces a different verdict than a 16 made up of strong workflow and weak knowledge readiness. The hard gates exist to encode that asymmetry.
How a score becomes a rationale
No score travels alone. Each dimension score is accompanied by a one-paragraph written rationale that references the specific evidence — interview quotes, document samples, system observations — that drove the number. The rationale is what makes the verdict defensible six months later, when leadership asks why a particular initiative was killed and the only person in the room with full context has moved on.
In practice this means every assessment produces, per initiative: six anchored scores, six paragraph-length rationales, a candidate verdict from the math, and a final verdict that incorporates judgement overlays. The trail is the deliverable.
§ 04
The hard gates
Three dimensions that can overrule the total.
Three of the six dimensions carry gate authority — when they score 0 or 1, they constrain or force the verdict regardless of how strong the other dimensions look. This is the mechanism that makes "Stop" structurally reachable. Without it, every assessment would average its way to "Proceed with Constraints," which is the failure mode of consultancy scoring frameworks generally.
Workflow Suitability ≤ 1
(strong knowledge, governance, etc. cannot rescue this)
→
Verdict locked to Convert to Deterministic Automation or Stop
Knowledge Readiness ≤ 1
(the model amplifies what it retrieves)
→
Verdict locked to Redesign or Stop
Operational Integration = 0
(structurally cannot be deployed)
→
Verdict forced to Stop
Operational Integration = 1
(can run beside, not inside)
→
Verdict locked to Stop or Keep Human-Assisted
Workflow ≥ 3 & Knowledge ≤ 1
(right fit, wrong substrate)
→
Verdict locked to Redesign — knowledge work must precede deployment
Trust & Governance ≤ 1 & regulated sector
(financial services, insurance, healthcare, energy)
→
Verdict locked to Stop or Redesign
§ 05
Turning scores into verdicts
Three stages, in order.
i.
Apply the gates
First, check the three hard-gate dimensions. If any gate condition fires, the verdict is constrained to a specific subset before the total is considered. Gates run before arithmetic — a strong total cannot overrule a failed gate.
ii.
Read the total
Within the verdict subset allowed by the gates, the total score across all six dimensions maps to a candidate verdict band. The bands are calibrated so that "Proceed" requires genuine strength across the board, not a high average with one fatal weakness.
iii.
Overlay judgement
The candidate verdict is reviewed against organisational readiness, time-to-value, vendor execution risk, and political reality. Judgment can lower a verdict (e.g. from Proceed to Proceed-with-Constraints) but cannot override a gate.
Score bands
20 – 24
Proceed
Production conditions already exist. Move. Strong scores across all dimensions, no gate failures.
15 – 19
Proceed with Constraints
Workable, but additional controls, narrower scope, or human approval required before deployment.
10 – 14
Redesign
Substantial gaps. Workflow, governance, or operating conditions must change before this is viable.
5 – 9
Keep Human-Assisted
AI supports decisions, does not execute them. Human judgement remains the unit of work. Often paired with low Failure-Path or Trust scores.
0 – 4
Stop
Poor workflow fit, weak economics, excessive operational risk. Or any gate at 0. Not every AI initiative should move forward.
Convert to Deterministic Automation sits outside the score bands — it is reached only via the Workflow Suitability gate. The math does not arbitrate this verdict; the workflow does.
§ 06
Connection-type taxonomy
Beyond verdict — shape of intervention.
"Proceed" alone is not a recommendation. Every Proceed and Proceed-with-Constraints verdict specifies the shape of intervention — how much autonomy, how much oversight, how much governance. The taxonomy below classifies the five intervention models in ascending order of autonomy and decreasing order of human involvement per decision.
Level 1
Deterministic Automation
Rule-based. No model in the production path. Outcomes fully predictable from inputs.
Reached via the Workflow gate when scoring shows AI is the wrong tool.
Autonomy: None
Level 2
Retrieval / Copilot
Model surfaces information, generates draft content, or summarises context. Human reads, decides, and acts.
Lowest-risk AI deployment shape. Suitable when Trust scores are still developing.
Autonomy: Low
Level 3
Decision Support
Model recommends a specific action with reasoning and confidence. Human approves and executes.
The native shape for "Keep Human-Assisted" verdicts in regulated workflows.
Autonomy: Medium
Level 4
Governed Agent (bounded)
Model acts within explicitly defined boundaries with full audit trail. Out-of-bounds cases escalate to humans.
Requires Failure-Path Design ≥ 3 and Governance ≥ 3.
Autonomy: High (bounded)
Level 5
Governed Agent (autonomous)
Model acts under policy enforcement, with oversight by sampling rather than approval per decision.
Requires near-perfect scoring across all dimensions and demonstrated bounded-agent track record.
Autonomy: High (sampled)
§ 07
The working method
How an engagement actually runs.
The framework is the discipline. The engagement is how the discipline meets the organisation. Four phases, in sequence, regardless of engagement tier.
Phase 01
Surface
Executive interviews, workflow walk-throughs, document inventory. Establish what is real versus what is documented. Identify the initiatives in scope and their sponsors.
Phase 02
Score
Apply the six-dimension rubric to each initiative. Capture evidence for every score. Produce candidate verdicts. Flag inconsistencies between stated workflow and observed workflow.
Phase 03
Test
Pressure-test candidate verdicts with sponsors and operators. Surface the political reality. Adjust judgement overlays. Verify no verdict has shifted on advocacy alone.
Phase 04
Deliver
Written memo per initiative with score, rationale, verdict, and recommended intervention level. Closing leadership session. Portfolio view. KPI model. Roadmap.
The most valuable phase is rarely Score. It is Test — where sponsors discover their own initiative scored a 1 on Knowledge Readiness, and the room has to decide whether to remediate or stop. The scoring is the instrument; the conversation it forces is the deliverable.
§ 08
Where the framework stops
What this is not.
Honest assessment requires honest limits. The framework is built for one specific question. Outside that question, it adds friction rather than clarity.
This is not a vendor selection framework. Scoring asks whether an initiative should proceed, not which platform should deliver it. Vendor execution-fit is a separate analysis, conducted only after a Proceed verdict.
This is not an AI strategy framework. Strategic questions — what AI capabilities the organisation should build, in what sequence, against what competitive positioning — are upstream of this work. The framework operates on initiatives that have already been proposed.
This is not a compliance audit. Trust & Governance scoring identifies whether governance is viable, not whether a specific implementation satisfies a specific regulator. Compliance audits remain compliance audits.
This is not a technical architecture review. Operational Integration scoring evaluates whether deployment is structurally feasible, not whether the chosen architecture is optimal. Architecture review is downstream of a Proceed verdict.
The framework also stops being useful where stakes are existential. For decisions whose failure could materially harm the firm — large-scale autonomous customer-facing deployments in regulated sectors, deployments touching safety-critical systems — the assessment is necessary but not sufficient. Those decisions require additional adversarial review by domain specialists. The framework will tell you whether to consider the deployment; it will not tell you that the deployment is safe.