Files
Ansible/scripts/docs/OPERATION_RESULTS.md
2026-09-22 19:23:17 +02:00

16 KiB

Purposeful operation results

AIM 3.3.0rc8. Publisher aim_output_v1; public aim_operation_result_v1. This is independent of play_task_host_v1 progress and target_outcome_summary_v1. The service, wire and event versions remain 1.0. It is not a raw-output/debug option.

Capability and catalog negotiation

capabilities.operation_results advertises publisher/result protocols, scopes, modes, limits and qualification. Each catalog operation has result: null or a declaration:

result:
  protocol: aim_output_v1
  schema: host_capabilities_v1
  scope: per_host
  required: true
  sensitivity: safe
  max_bytes_per_host: 1048576
  schema_file: host_capabilities_v1.yml

schema_file is a basename under the catalog's schemas/ directory. Core resolves it into data_schema for public list_playbooks and PreparedRun.result_contract output. Clients never pass a schema file or payload in an execute request. Schema/source bytes participate in the review revision. A changed schema requires preparation/review again. No hard-coded playbook dispatch exists in Core.

Publisher convention

- name: AIM | Publish operation result
  ansible.builtin.set_stats:
    per_host: true
    aggregate: false
    data:
      aim_output:
        protocol: aim_output_v1
        schema: host_capabilities_v1
        data:
          is_dc: '{{ is_dc | default(false) | bool }}'
          is_dhcp_server: '{{ is_dhcp_server | default(false) | bool }}'
          is_hyperv_host: '{{ is_hyperv_host | default(false) | bool }}'
          has_veeam_vbr: '{{ has_veeam_vbr | default(false) | bool }}'
          has_veeam_vbo: '{{ has_veeam_vbo | default(false) | bool }}'
          has_veeam_em: '{{ has_veeam_em | default(false) | bool }}'
          is_unifi_controller: '{{ is_unifi_controller | default(false) | bool }}'
          is_unifi_os_server: '{{ is_unifi_os_server | default(false) | bool }}'

Use exactly one publication per requested host/run. Do not loop publishers, publish preliminary values under the same key, or aggregate dictionaries accidentally. For a global declaration use per_host: false and exactly one global publication (beware run_once with serial batches). The bundled schemas are all per-host.

The callback observes only an actual set_stats result with a matching declaration, false aggregation and matching scope. It records sensitivity provenance and reconciles that value with final custom statistics. An unrelated custom stat, a debug task named "Publish", or a forged final statistic is not an output source. no_log publishers are withheld; custom stats alone cannot erase their provenance. Report from a reviewed normalizer with selected non-secret fields, not arbitrary registered result dictionaries.

Publishing set_stats is controller-side and adds no remote package/module dependency. Native console rendering remains unchanged; use the existing debug summaries for CLI operators. External clients obtain reports only at finalization, not through debug text.

Final public object

Every RunResult adds operation_result, null for an undeclared operation. Otherwise:

{
  "protocol": "aim_operation_result_v1",
  "schema": "host_capabilities_v1",
  "scope": "per_host",
  "required": true,
  "complete": true,
  "check_mode": false,
  "hosts": {
    "host01.example": {
      "schema": "host_capabilities_v1",
      "status": "available",
      "data": {"is_dc": false, "is_dhcp_server": false, "is_hyperv_host": false,
        "has_veeam_vbr": false, "has_veeam_vbo": false, "has_veeam_em": false,
        "is_unifi_controller": false, "is_unifi_os_server": false},
      "error": null
    }
  },
  "global": null
}

For global scope, hosts is empty and global holds the same entry shape. Schema and scope come from the catalog, not client input. Host keys come from reviewed targets. The existing final result event contains the same object as the final response. There is no separate live operation-data event; read final authoritative results in either summary or detail mode.

Entry status:

Status Meaning
available Complete final data received, schema/limit/safety validation passed
missing Final accounting arrived but this report slot was not published
withheld Sensitive publisher or known supplied secret matched output
invalid Schema, provenance or size validation rejected the publication
not_started Execution did not start; there is no operation data
indeterminate Work may have started but no complete final report stream exists

Unavailable entries have data: null; never replace them with {}/zero/false in a UI. complete describes report availability, not execution success or payload-specific completeness. A patch report has its own data.complete for update evidence. A failed service-start operation can have an available, complete report listing failed services.

If native exit is nonzero, preserve the native/Core failure and retain any valid reports from the final stats. If native execution exits 0 but required output is missing, withheld, invalid or incomplete, Core returns failed/result_validation with native exit_code 0. Native target outcomes still describe their Ansible stats. This is an intentional additional report-contract check, not a redefinition of native stats. Optional missing reports do not fail an otherwise successful run; invalid submitted reports do. Never replay automatically after result-validation failure.

Generic schema subset and safety

The schema validator is standard-library-only, shared by callback and Core. Supported JSON Schema keywords: type (including nullable type arrays), properties, required, boolean additionalProperties, items, enum, maxItems, maxLength, minimum, maximum, and documentation description. $ref, arbitrary validators and executable schema extensions are not supported. Objects are closed unless explicitly declared open. The only shipped open subtree is already parsed/redacted Checkmk sections.

Limits: maximum depth 20, 200,000 nodes per value, 20,000 items per collection, 8,192 characters per string, 1 MiB encoded envelope per host/global slot and 16 MiB per run. Catalogs can lower the per-host byte cap. Data is chunked privately into <=2,048-byte pieces; malformed sequences, duplicate publications, torn frames and oversize data are rejected, never silently truncated. NaN/Infinity, binary objects, cycles and unknown closed-object properties are rejected. Schemas and output labels must be static public controller source.

Obvious secret field names are disallowed in schemas; redacted dynamic config keys can only hold the exact [REDACTED] marker. Known supplied credentials matching public strings are withheld again at the Core boundary. These measures are defense in depth, not a general secret scanner. Config/other strings literally containing an undisclosed secret remain an author/operator responsibility. Never publish environments, invocation, exception objects, arbitrary debug values, commands or private keys.

checkmk_user_config_v1 preserves parsed sections but redacts recognized password/token/ passphrase/credential/community fields, command/argument/environment bodies, MRPE commands, URL credentials and obvious inline secrets. redacted_paths explains what was withheld. Comments/formatting are not YAML data. The operation may read only basename check_mk.user.yml, <=512 KiB, from a trusted configured path. It never modifies the file. Unknown settings remain visible unless filtered; do not advertise this as guaranteed secret-free content. No general arbitrary-file reader is added.

Current schemas and exact semantic boundaries

Normative field shapes: playbooks/schemas/<schema>.yml (also in catalog metadata).

Schema Data meaning
host_capabilities_v1 Eight booleans, source detection facts; no inventory membership edits
filesystem_usage_v1 Selected operational mounts/volumes; byte quantities and nullable observations; Windows uses attached storage volumes from community.windows.win_disk_facts and excludes mapped/network drives
event_log_export_v1 Channel names, age and target file paths; no log contents or controller download; check mode has no exported paths
service_start_summary_v1 Before stopped services, policy eligibility, attempts, exclusions, actual after observations and fixed per-attempt failures
patch_summary_v1 Per-package net version-set changes on Debian/RedHat, or Windows update IDs/titles/KBs with installed flags; never arbitrary manager dictionaries
managed_cleanup_preview_v1 Managed filename candidates and observed deletions; no unknown-file purge
checkmk_user_config_v1 Parsed/redacted sections and file metadata, not raw YAML/debug output
checkmk_agent_state_v1 Installed status/version from registry/package query, actual service states, package/config/check changes
checkmk_agent_config_v1 Named section/config/check change actions; no raw before/after values or file diffs

Services

initially_stopped is the observation before action. eligible applies auto/delayed-start and configured include/exclude policy. attempted excludes skipped and check-mode actions. newly_running is the observed stopped-to-started transition, including a concurrent external start; started_count counts only attempted services now running. It is not proof that the start return alone kept a service alive. still_stopped includes excluded services; unobserved means missing after observations. Non-running attempted services appear in failed_to_start with name, fixed reason/message and nullable numeric native code. Failed reasons include dependency_failed, permission_denied, service_disabled, start_timeout, service_not_found, logon_failed, start_failed, not_running_after_start and state_unavailable. Localized unrecognized errors use generic start_failed; no raw text. An unavailable post-query fails the run instead of inventing a successful observation.

Packages and Checkmk

Linux queries native package databases before/after the update/reboot policy, preserving architecture and parallel installed version sets. updates lists updated/installed/ removed rows and separate counters. Old/new arrays are empty where the package did not exist. Snapshots cannot identify an intermediate reinstall with unchanged final versions or attribute concurrent external package changes; do not run another package manager concurrently. Check mode reports no completed changes, not predicted upgrade versions. No silently shortened list: exceeding limits becomes explicit validation failure.

Windows reports native updates marked installed, pending/failed records and numeric failure codes. A mismatch between native installed count and detailed records marks data.complete: false, rather than fabricating records. Supported scope remains Windows Update under configured categories, not every third-party package installer. Reboot facts retain their observed meaning; check mode never claims a performed reboot.

Checkmk versions can be null (unavailable/ambiguous); do not fall back to the filename. Config changes are section/file-level, not field-level raw diffs. deployed_checks means selected file-copy tasks completed; changed_checks identifies changed copy results. Check-mode changes are predictions; mode is always explicit. Unknown files are neither counted nor scanned: policy is untouched_not_enumerated, not an invented preserved count.

Errors and consumer behavior

Report errors: operation_result_missing, operation_result_invalid, operation_result_withheld, operation_result_limit, operation_result_incomplete. Messages contain no payload/exception content. Preserve final Core/native status; show report issues separately; no credential retry or task replay is implied.

Treat this candidate as an implementation awaiting native 2.19.11/managed-host acceptance. Upstream interfaces consulted: Ansible 2.19.11 plugins/action/set_stats.py, executor/stats.py, and callback result contracts; ansible.windows >=3.8.0,<4.0.0 win_updates and win_service_info return definitions. Consultation is not native testing.

Patch wave and reboot semantics (3.3.0rc8)

patch_summary_v1 reports package/update evidence independently from reboot policy. Linux retains the rc2 package-manager transaction model. Windows delegates each selected patch wave to one native ansible.windows.win_updates state: installed invocation with reboot: false. The collection and Windows Update Agent own update ordering/coordination inside that wave; AIM does not schedule updates individually.

The returned update dictionary is normalized into reviewed installed/failed records. A module-level failure with no per-update records becomes a bounded generic failed-wave record; raw failure text, exceptions and arbitrary module dictionaries are not published.

A newly required reboot with os_patching_reboot: false does not invalidate successful installation: reboot_required/reboot_required_after are true and reboot_deferred is true. The run stops. On a later run, a detected pre-existing reboot with automatic reboot still disabled publishes evidence: preflight_reboot_state, complete: false, blocked_reason: preexisting_reboot_required, then fails before new patch work starts.

When automatic reboot is enabled, AIM waits for the native update wave to return before performing its reviewed message/delay/reboot. By default os_patching_rescan_after_reboot: false; any AIM-performed reboot ends the run with continuation_required: true and remaining_updates_known: false. This includes a reboot that was already required before patching began. A new operator-approved run owns discovery of the next patch state.

If the operator explicitly sets post-reboot continuation true, AIM may invoke another native Windows update wave after reboot. Continuation is bounded to 12 wave invocations per run. A failed wave is never automatically replayed or used as authority to continue.

When a Windows wave completes without reboot, AIM performs one final read-only search. This does not install newly applicable updates. remaining_updates_known is then true and pending contains that final observation; continuation_required indicates whether another operator-approved run has applicable work.

patch_cycles counts native Windows install-wave invocations entered in the run and rescan_after_reboot echoes the reviewed option. Those fields are Windows-specific and are optional in the closed schema so existing Linux result shape remains stable.

Windows failed_updates contains only bounded safe fields: update identity/title, normalized unsigned HRESULT and hex form, a fixed reason and fixed message. Known reasons include operation_in_progress, install_not_allowed, not_applicable, exclusive_install_conflict, self_update_in_progress, no_connection, and timeout. Raw Windows Update failure strings are not published. 0x80240016 maps to install_not_allowed; it does not prove a reboot was pending because Windows also uses that HRESULT while another installation is active. Only the independent preflight can report preexisting_reboot_required.

reboot_required_before records the pre-run observation. reboot_performed is true when AIM completed an approved reboot. reboot_required_after and the legacy reboot_required field describe the final pending state represented by the report. reboot_delay_minutes echoes the reviewed delay. The user-facing reboot message itself is not copied into the structured result.

Windows, Debian and RedHat use platform-native reboot-required signals; RedHat preflight is conservative when needs-restarting is not already present. Report fields never imply that every possible vendor-specific reboot indicator was discovered.