Files
Ansible/scripts/docs/OPERATION_RESULTS.md
T
2026-09-22 19:23:17 +02:00

277 lines
16 KiB
Markdown

# Purposeful operation results
**AIM 3.3.0rc8. Publisher `aim_output_v1`; public `aim_operation_result_v1`.**
This is independent of `play_task_host_v1` progress and `target_outcome_summary_v1`.
The service, wire and event versions remain 1.0. It is not a raw-output/debug option.
## Capability and catalog negotiation
`capabilities.operation_results` advertises publisher/result protocols, scopes, modes,
limits and qualification. Each catalog operation has `result: null` or a declaration:
```yaml
result:
protocol: aim_output_v1
schema: host_capabilities_v1
scope: per_host
required: true
sensitivity: safe
max_bytes_per_host: 1048576
schema_file: host_capabilities_v1.yml
```
`schema_file` is a basename under the catalog's `schemas/` directory. Core resolves it
into `data_schema` for public `list_playbooks` and `PreparedRun.result_contract` output.
Clients never pass a schema file or payload in an execute request. Schema/source bytes
participate in the review revision. A changed schema requires preparation/review again.
No hard-coded playbook dispatch exists in Core.
## Publisher convention
```yaml
- name: AIM | Publish operation result
ansible.builtin.set_stats:
per_host: true
aggregate: false
data:
aim_output:
protocol: aim_output_v1
schema: host_capabilities_v1
data:
is_dc: '{{ is_dc | default(false) | bool }}'
is_dhcp_server: '{{ is_dhcp_server | default(false) | bool }}'
is_hyperv_host: '{{ is_hyperv_host | default(false) | bool }}'
has_veeam_vbr: '{{ has_veeam_vbr | default(false) | bool }}'
has_veeam_vbo: '{{ has_veeam_vbo | default(false) | bool }}'
has_veeam_em: '{{ has_veeam_em | default(false) | bool }}'
is_unifi_controller: '{{ is_unifi_controller | default(false) | bool }}'
is_unifi_os_server: '{{ is_unifi_os_server | default(false) | bool }}'
```
Use exactly one publication per requested host/run. Do not loop publishers, publish
preliminary values under the same key, or aggregate dictionaries accidentally. For a
global declaration use `per_host: false` and exactly one global publication (beware
run_once with serial batches). The bundled schemas are all per-host.
The callback observes only an actual set_stats result with a matching declaration,
false aggregation and matching scope. It records sensitivity provenance and reconciles
that value with final custom statistics. An unrelated custom stat, a debug task named
"Publish", or a forged final statistic is not an output source. `no_log` publishers
are withheld; custom stats alone cannot erase their provenance. Report from a reviewed
normalizer with selected non-secret fields, not arbitrary registered result dictionaries.
Publishing set_stats is controller-side and adds no remote package/module dependency.
Native console rendering remains unchanged; use the existing debug summaries for CLI
operators. External clients obtain reports only at finalization, not through debug text.
## Final public object
Every RunResult adds `operation_result`, null for an undeclared operation. Otherwise:
```json
{
"protocol": "aim_operation_result_v1",
"schema": "host_capabilities_v1",
"scope": "per_host",
"required": true,
"complete": true,
"check_mode": false,
"hosts": {
"host01.example": {
"schema": "host_capabilities_v1",
"status": "available",
"data": {"is_dc": false, "is_dhcp_server": false, "is_hyperv_host": false,
"has_veeam_vbr": false, "has_veeam_vbo": false, "has_veeam_em": false,
"is_unifi_controller": false, "is_unifi_os_server": false},
"error": null
}
},
"global": null
}
```
For global scope, `hosts` is empty and `global` holds the same entry shape. Schema and
scope come from the catalog, not client input. Host keys come from reviewed targets.
The existing final result event contains the same object as the final response. There
is no separate live operation-data event; read final authoritative results in either
summary or detail mode.
Entry status:
| Status | Meaning |
|---|---|
| available | Complete final data received, schema/limit/safety validation passed |
| missing | Final accounting arrived but this report slot was not published |
| withheld | Sensitive publisher or known supplied secret matched output |
| invalid | Schema, provenance or size validation rejected the publication |
| not_started | Execution did not start; there is no operation data |
| indeterminate | Work may have started but no complete final report stream exists |
Unavailable entries have `data: null`; never replace them with `{}`/zero/false in a UI.
`complete` describes report availability, not execution success or payload-specific
completeness. A patch report has its own `data.complete` for update evidence. A failed
service-start operation can have an available, complete report listing failed services.
If native exit is nonzero, preserve the native/Core failure and retain any valid reports
from the final stats. If native execution exits 0 but required output is missing,
withheld, invalid or incomplete, Core returns failed/result_validation with native
exit_code 0. Native target outcomes still describe their Ansible stats. This is an
intentional additional report-contract check, not a redefinition of native stats.
Optional missing reports do not fail an otherwise successful run; invalid submitted
reports do. Never replay automatically after result-validation failure.
## Generic schema subset and safety
The schema validator is standard-library-only, shared by callback and Core. Supported
JSON Schema keywords: `type` (including nullable type arrays), `properties`, `required`,
boolean `additionalProperties`, `items`, `enum`, `maxItems`, `maxLength`, `minimum`,
`maximum`, and documentation `description`. `$ref`, arbitrary validators and executable
schema extensions are not supported. Objects are closed unless explicitly declared open.
The only shipped open subtree is already parsed/redacted Checkmk `sections`.
Limits: maximum depth 20, 200,000 nodes per value, 20,000 items per collection, 8,192
characters per string, 1 MiB encoded envelope per host/global slot and 16 MiB per run.
Catalogs can lower the per-host byte cap. Data is chunked privately into <=2,048-byte
pieces; malformed sequences, duplicate publications, torn frames and oversize data are
rejected, never silently truncated. NaN/Infinity, binary objects, cycles and unknown
closed-object properties are rejected. Schemas and output labels must be static public
controller source.
Obvious secret field names are disallowed in schemas; redacted dynamic config keys
can only hold the exact `[REDACTED]` marker. Known supplied credentials matching public
strings are withheld again at the Core boundary. These measures are defense in depth,
not a general secret scanner. Config/other strings literally containing an undisclosed
secret remain an author/operator responsibility. Never publish environments, invocation,
exception objects, arbitrary debug values, commands or private keys.
`checkmk_user_config_v1` preserves parsed sections but redacts recognized password/token/
passphrase/credential/community fields, command/argument/environment bodies, MRPE
commands, URL credentials and obvious inline secrets. `redacted_paths` explains what
was withheld. Comments/formatting are not YAML data. The operation may read only basename
`check_mk.user.yml`, <=512 KiB, from a trusted configured path. It never modifies the file.
Unknown settings remain visible unless filtered; do not advertise this as guaranteed
secret-free content. No general arbitrary-file reader is added.
## Current schemas and exact semantic boundaries
Normative field shapes: `playbooks/schemas/<schema>.yml` (also in catalog metadata).
| Schema | Data meaning |
|---|---|
| host_capabilities_v1 | Eight booleans, source detection facts; no inventory membership edits |
| filesystem_usage_v1 | Selected operational mounts/volumes; byte quantities and nullable observations; Windows uses attached storage volumes from community.windows.win_disk_facts and excludes mapped/network drives |
| event_log_export_v1 | Channel names, age and target file paths; no log contents or controller download; check mode has no exported paths |
| service_start_summary_v1 | Before stopped services, policy eligibility, attempts, exclusions, actual after observations and fixed per-attempt failures |
| patch_summary_v1 | Per-package net version-set changes on Debian/RedHat, or Windows update IDs/titles/KBs with installed flags; never arbitrary manager dictionaries |
| managed_cleanup_preview_v1 | Managed filename candidates and observed deletions; no unknown-file purge |
| checkmk_user_config_v1 | Parsed/redacted sections and file metadata, not raw YAML/debug output |
| checkmk_agent_state_v1 | Installed status/version from registry/package query, actual service states, package/config/check changes |
| checkmk_agent_config_v1 | Named section/config/check change actions; no raw before/after values or file diffs |
### Services
`initially_stopped` is the observation before action. `eligible` applies auto/delayed-start
and configured include/exclude policy. `attempted` excludes skipped and check-mode actions.
`newly_running` is the observed stopped-to-started transition, including a concurrent
external start; `started_count` counts only attempted services now running. It is not
proof that the start return alone kept a service alive. `still_stopped` includes excluded
services; `unobserved` means missing after observations. Non-running attempted services
appear in `failed_to_start` with name, fixed reason/message and nullable numeric native
code. Failed reasons include dependency_failed, permission_denied, service_disabled,
start_timeout, service_not_found, logon_failed, start_failed, not_running_after_start and
state_unavailable. Localized unrecognized errors use generic start_failed; no raw text.
An unavailable post-query fails the run instead of inventing a successful observation.
### Packages and Checkmk
Linux queries native package databases before/after the update/reboot policy, preserving
architecture and parallel installed version sets. `updates` lists updated/installed/
removed rows and separate counters. Old/new arrays are empty where the package did not
exist. Snapshots cannot identify an intermediate reinstall with unchanged final versions
or attribute concurrent external package changes; do not run another package manager
concurrently. Check mode reports no completed changes, not predicted upgrade versions.
No silently shortened list: exceeding limits becomes explicit validation failure.
Windows reports native updates marked installed, pending/failed records and numeric
failure codes. A mismatch between native installed count and detailed records marks
`data.complete: false`, rather than fabricating records. Supported scope remains Windows
Update under configured categories, not every third-party package installer. Reboot
facts retain their observed meaning; check mode never claims a performed reboot.
Checkmk versions can be null (unavailable/ambiguous); do not fall back to the filename.
Config changes are section/file-level, not field-level raw diffs. `deployed_checks` means
selected file-copy tasks completed; `changed_checks` identifies changed copy results.
Check-mode changes are predictions; mode is always explicit. Unknown files are neither
counted nor scanned: policy is `untouched_not_enumerated`, not an invented preserved count.
## Errors and consumer behavior
Report errors: operation_result_missing, operation_result_invalid,
operation_result_withheld, operation_result_limit, operation_result_incomplete.
Messages contain no payload/exception content. Preserve final Core/native status;
show report issues separately; no credential retry or task replay is implied.
Treat this candidate as an implementation awaiting native 2.19.11/managed-host acceptance.
Upstream interfaces consulted: Ansible 2.19.11 `plugins/action/set_stats.py`,
`executor/stats.py`, and callback result contracts; ansible.windows >=3.8.0,<4.0.0 win_updates and
win_service_info return definitions. Consultation is not native testing.
## Patch wave and reboot semantics (3.3.0rc8)
`patch_summary_v1` reports package/update evidence independently from reboot policy. Linux
retains the rc2 package-manager transaction model. Windows delegates each selected patch
wave to one native `ansible.windows.win_updates` `state: installed` invocation with
`reboot: false`. The collection and Windows Update Agent own update ordering/coordination
inside that wave; AIM does not schedule updates individually.
The returned update dictionary is normalized into reviewed installed/failed records. A
module-level failure with no per-update records becomes a bounded generic failed-wave record;
raw failure text, exceptions and arbitrary module dictionaries are not published.
A newly required reboot with `os_patching_reboot: false` does not invalidate successful
installation: `reboot_required`/`reboot_required_after` are true and `reboot_deferred` is
true. The run stops. On a later run, a detected pre-existing reboot with automatic reboot
still disabled publishes `evidence: preflight_reboot_state`, `complete: false`,
`blocked_reason: preexisting_reboot_required`, then fails before new patch work starts.
When automatic reboot is enabled, AIM waits for the native update wave to return before
performing its reviewed message/delay/reboot. By default
`os_patching_rescan_after_reboot: false`; any AIM-performed reboot ends the run with
`continuation_required: true` and `remaining_updates_known: false`. This includes a reboot
that was already required before patching began. A new operator-approved run owns discovery
of the next patch state.
If the operator explicitly sets post-reboot continuation true, AIM may invoke another native
Windows update wave after reboot. Continuation is bounded to 12 wave invocations per run. A
failed wave is never automatically replayed or used as authority to continue.
When a Windows wave completes without reboot, AIM performs one final read-only search. This
does not install newly applicable updates. `remaining_updates_known` is then true and
`pending` contains that final observation; `continuation_required` indicates whether another
operator-approved run has applicable work.
`patch_cycles` counts native Windows install-wave invocations entered in the run and
`rescan_after_reboot` echoes the reviewed option. Those fields are Windows-specific and are
optional in the closed schema so existing Linux result shape remains stable.
Windows `failed_updates` contains only bounded safe fields: update identity/title, normalized
unsigned HRESULT and hex form, a fixed reason and fixed message. Known reasons include
`operation_in_progress`, `install_not_allowed`, `not_applicable`,
`exclusive_install_conflict`, `self_update_in_progress`, `no_connection`, and `timeout`.
Raw Windows Update failure strings are not published. `0x80240016` maps to
`install_not_allowed`; it does not prove a reboot was pending because Windows also uses that
HRESULT while another installation is active. Only the independent preflight can report
`preexisting_reboot_required`.
`reboot_required_before` records the pre-run observation. `reboot_performed` is true when
AIM completed an approved reboot. `reboot_required_after` and the legacy
`reboot_required` field describe the final pending state represented by the report.
`reboot_delay_minutes` echoes the reviewed delay. The user-facing reboot message itself is
not copied into the structured result.
Windows, Debian and RedHat use platform-native reboot-required signals; RedHat preflight
is conservative when `needs-restarting` is not already present. Report fields never imply
that every possible vendor-specific reboot indicator was discovered.