Back to Field Notes
Project Controls & DeliveryAI project controlsConstruction schedulingSchedule quality

AI Does Not Fix a Weak Schedule: Build an Evidence Gate for Every Review

A practical control plan for Iranian projects that lets AI triage schedule anomalies without changing the baseline, inventing delay causes, or replacing contractual judgment.

By OlbrichCo Technical OfficePublished 10 min read
A cobalt optical inspector on a steel arm examines one broken linkage in a chain of concrete schedule blocks
A cobalt optical inspector on a steel arm examines one broken linkage in a chain of concrete schedule blocks

A language model cannot repair a weak schedule

The GAO Schedule Assessment Guide treats a reliable schedule as a model of the work and time needed to reach project milestones, built through ten connected practices and maintained against an approved plan. DOE guidance similarly begins analysis by checking data validity and schedule health before examining variance, trend, or forecast. An eloquent narrative cannot compensate for missing scope, broken logic, uncontrolled calendars, or unreliable status. [1][2]

OlbrichCo analysis: use AI as a reviewer of a controlled programme, never as the source of programme truth. If the input cannot reproduce the approved baseline, current update, data date, calendars, progress rules, and change history, stop the AI review. Fix the schedule-management process first; otherwise automation merely produces a faster, more persuasive version of the same uncertainty.

Define the decision and its consequence before choosing the tool

ISO/IEC 42001 frames AI use as a managed lifecycle with policies, objectives, risk treatment, and continual improvement. ISO/IEC 42005 adds impact assessment throughout that lifecycle, while the UK Government AI Playbook calls for meaningful human control at the stages where decisions carry greater risk. These sources support a use-case boundary, not blanket permission to automate project controls. [5][6][7]

Classify the intended output before procurement. Low-consequence assistance may group duplicate comments or draft questions. Medium-consequence use may rank logic anomalies for a planner. High-consequence use includes changing the baseline, accepting progress, forecasting a contractual milestone, attributing delay, approving an extension of time, or supporting payment. Keep that final class outside autonomous action and name the accountable human decision-maker.

  • State the user, decision, input set, allowed output, prohibited action, and review owner.
  • Define what a false negative and false positive would do to safety, cost, time, and rights.
  • Choose deterministic software for calculations and AI only where language or pattern triage adds value.
  • Set an escalation route for unsupported claims, conflicting records, and uncertain activity identity.
  • Do not let a vendor demonstration define the project's risk tolerance.

Freeze a review packet that every finding can cite

GAO and DOE guidance depend on an integrated, controlled schedule and valid status data. NIST's Generative AI Profile separately recommends establishing data origin and content lineage, verifying sources and citations in outputs, and documenting limits beyond the tested setting. Together they point to one prerequisite: the review must be reproducible from a named evidence set. [1][2][4]

Export a read-only review packet rather than connecting an opaque assistant directly to the live scheduling file. Include the native file, neutral tabular extract, approved baseline reference, update identifier, data date, calendars, coding dictionary, progress basis, change register, narrative, and relevant contract milestones. Hash or otherwise identify the packet; record tool and model version, configuration, prompt, retrieval set, run time, and reviewer.

  • Preserve Persian and English activity names plus a stable activity ID; never join records by translated text alone.
  • Store Gregorian dates unambiguously and document any Jalali display conversion at the boundary.
  • Resolve workweek, holiday, shift, weather, procurement, and commissioning calendars before analysis.
  • Keep site updates usable offline; synchronize only controlled packets when connectivity permits.
  • Remove or mask personal, commercial, and security-sensitive data not required for the stated review.

Put deterministic schedule tests ahead of probabilistic review

GAO's schedule practices address complete activities, sequencing, resources, duration, traceable status, critical path, float, risk analysis, and controlled updates. DOE's analysis sequence checks validity and health before variance and forecast. NIST advises empirical evaluation of model claims and warns against extrapolating performance from narrow or anecdotal tests. Schedule arithmetic should therefore remain in a transparent scheduling engine or test script. [1][2][4]

Run a documented rule set first and give the AI only the resulting exceptions plus the evidence needed to explain them. A planner should be able to reproduce every duration, float, variance, and logic trace without the model. The AI may then cluster related exceptions, compare the update narrative with recorded changes, and draft targeted questions; it must not invent missing relationships or silently recalculate contractual dates.

  • Missing predecessors or successors, open ends, loops, redundant logic, and excessive constraints.
  • Actual or forecast dates inconsistent with the data date and approved progress rules.
  • Calendar, lag, duration, float, and critical-path changes that exceed agreed review thresholds.
  • Out-of-sequence work, retained logic choices, and unexplained changes to remaining duration.
  • Milestones, procurement, approvals, access, interfaces, testing, and handover absent from the integrated logic.
  • Baseline-to-current additions, deletions, recoding, and logic edits without an approved change reference.

Constrain the model to evidence-shaped questions

NIST identifies confabulation and information-integrity risk in generative AI, recommends reviewing sources and citations, and calls for monitoring human overrides. The UK playbook requires humans to validate high-risk decisions and recommends quantitative testing and validation within change control. A fluent answer is therefore not an acceptance criterion. [4][7]

Require a fixed output schema. Every finding should identify activity IDs, evidence rows, the deterministic rule or record conflict, why it matters, plausible explanations labelled as hypotheses, a question for the responsible party, and the reviewer disposition. Reject any finding that cannot cite the packet. The safe result is often a better question, not a predicted completion date.

  • Fact: what the controlled packet explicitly contains.
  • Calculation: what the deterministic engine reproduces.
  • Inference: a labelled interpretation that still needs domain review.
  • Unknown: missing evidence, ambiguous coding, or conflicting records.
  • Disposition: accept, reject, investigate, correct data, or escalate—with a named reviewer and date.

Commission the review on past updates before live use

NIST recommends pre-deployment testing in conditions similar to use, documenting generalization limits, verifying citations, and monitoring performance after release. ISO/IEC 42001 and 42005 treat assessment and improvement as lifecycle work; the UK playbook also requires managed releases, quantitative validation, ongoing monitoring, and a route to revert changes. A one-time demo is not commissioning. [4][5][6][7]

Build a blind test set from several closed schedule updates: clean cases, known logic defects, bilingual naming, calendar changes, revised coding, disputed progress, and incomplete narratives. Have experienced planners establish the reference findings without seeing the AI output. Approve a narrow use only when results are stable enough for its consequence class; repeat the test after model, prompt, data pipeline, calendar, or schedule-software changes.

  • Recall of known material anomalies and rate of material anomalies missed.
  • False-positive rate and reviewer time spent clearing noise.
  • Percentage of findings with valid activity IDs, source rows, and reproducible calculations.
  • Unsupported-cause, invented-date, and citation-error count—target zero before live review.
  • Human override rate, disagreement pattern, run-to-run stability, and time to close each finding.

Pilot one update cycle and keep contractual decisions human

The NIST AI RMF organizes risk work across governance, context, measurement, and management. ISO/IEC 42001 requires an organizational system rather than a tool-only control, and the UK playbook calls for clear oversight, roles, escalation, monitoring, and meaningful intervention. Those controls matter most when AI output may influence people, payment, programme commitments, or legal positions. [3][5][7]

Pilot one contractor update in parallel with the existing review. Do not write back to the live file. Compare the AI-assisted exception list with the normal planner review, close each finding through the same contractual communication route, and hold a go/no-go review. Expand only if evidence quality improves without hiding uncertainty or weakening independent professional judgment.

Final acceptance of progress, baseline changes, delay causation, concurrency, mitigation, milestone forecasts, extensions of time, and payment remains a project and contract decision. Governing Iranian requirements, the signed contract, approved schedules, contemporaneous records, actual site conditions, and responsible planning, engineering, commercial, legal, and information-security review control. This workflow is an assurance gate, not proof of entitlement or a performance guarantee.

  • Zero AI-originated changes written directly into the approved or current schedule.
  • Every material finding traceable to a frozen packet, reproducible rule, and named reviewer.
  • Every model or configuration change triggers proportionate regression testing and release approval.
  • Every override, failure, incident, and user complaint enters the improvement register.
  • A documented off-switch and manual review path remain available for every cycle.

Sources & further reading

These primary sources support the claims and implementation frameworks used in this field note.

  1. 1. GAO-16-89G — Schedule Assessment Guide: Best Practices for Project Schedules

    U.S. Government Accountability Office

  2. 2. EVMS Implementation Guidance — Planning, Scheduling, Data Validity, and Analysis

    U.S. Department of Energy — Office of Project Management

  3. 3. Artificial Intelligence Risk Management Framework (AI RMF 1.0)

    National Institute of Standards and Technology

  4. 4. NIST AI 600-1 — Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile

    National Institute of Standards and Technology

  5. 5. ISO/IEC 42001:2023 — Information technology: Artificial intelligence management system

    International Organization for Standardization

  6. 6. ISO/IEC 42005:2025 — Artificial intelligence system impact assessment

    International Organization for Standardization

  7. 7. Artificial Intelligence Playbook for the UK Government

    UK Government Digital Service

Sources were checked on 29 August 2026. The foreign standards and public guides describe schedule-management and AI-risk practices in their own contexts; they do not create Iranian legal obligations. The baseline, calendars, progress rules, data security and hosting, approval authority, delay assessment, and time entitlement must follow governing requirements, the signed contract, approved project procedures, and responsible planning, contract, legal, information-security, and delivery review.