277 lines
16 KiB
Markdown
277 lines
16 KiB
Markdown
# Purposeful operation results
|
|
|
|
**AIM 3.3.0rc8. Publisher `aim_output_v1`; public `aim_operation_result_v1`.**
|
|
This is independent of `play_task_host_v1` progress and `target_outcome_summary_v1`.
|
|
The service, wire and event versions remain 1.0. It is not a raw-output/debug option.
|
|
|
|
## Capability and catalog negotiation
|
|
|
|
`capabilities.operation_results` advertises publisher/result protocols, scopes, modes,
|
|
limits and qualification. Each catalog operation has `result: null` or a declaration:
|
|
|
|
```yaml
|
|
result:
|
|
protocol: aim_output_v1
|
|
schema: host_capabilities_v1
|
|
scope: per_host
|
|
required: true
|
|
sensitivity: safe
|
|
max_bytes_per_host: 1048576
|
|
schema_file: host_capabilities_v1.yml
|
|
```
|
|
|
|
`schema_file` is a basename under the catalog's `schemas/` directory. Core resolves it
|
|
into `data_schema` for public `list_playbooks` and `PreparedRun.result_contract` output.
|
|
Clients never pass a schema file or payload in an execute request. Schema/source bytes
|
|
participate in the review revision. A changed schema requires preparation/review again.
|
|
No hard-coded playbook dispatch exists in Core.
|
|
|
|
## Publisher convention
|
|
|
|
```yaml
|
|
- name: AIM | Publish operation result
|
|
ansible.builtin.set_stats:
|
|
per_host: true
|
|
aggregate: false
|
|
data:
|
|
aim_output:
|
|
protocol: aim_output_v1
|
|
schema: host_capabilities_v1
|
|
data:
|
|
is_dc: '{{ is_dc | default(false) | bool }}'
|
|
is_dhcp_server: '{{ is_dhcp_server | default(false) | bool }}'
|
|
is_hyperv_host: '{{ is_hyperv_host | default(false) | bool }}'
|
|
has_veeam_vbr: '{{ has_veeam_vbr | default(false) | bool }}'
|
|
has_veeam_vbo: '{{ has_veeam_vbo | default(false) | bool }}'
|
|
has_veeam_em: '{{ has_veeam_em | default(false) | bool }}'
|
|
is_unifi_controller: '{{ is_unifi_controller | default(false) | bool }}'
|
|
is_unifi_os_server: '{{ is_unifi_os_server | default(false) | bool }}'
|
|
```
|
|
|
|
Use exactly one publication per requested host/run. Do not loop publishers, publish
|
|
preliminary values under the same key, or aggregate dictionaries accidentally. For a
|
|
global declaration use `per_host: false` and exactly one global publication (beware
|
|
run_once with serial batches). The bundled schemas are all per-host.
|
|
|
|
The callback observes only an actual set_stats result with a matching declaration,
|
|
false aggregation and matching scope. It records sensitivity provenance and reconciles
|
|
that value with final custom statistics. An unrelated custom stat, a debug task named
|
|
"Publish", or a forged final statistic is not an output source. `no_log` publishers
|
|
are withheld; custom stats alone cannot erase their provenance. Report from a reviewed
|
|
normalizer with selected non-secret fields, not arbitrary registered result dictionaries.
|
|
|
|
Publishing set_stats is controller-side and adds no remote package/module dependency.
|
|
Native console rendering remains unchanged; use the existing debug summaries for CLI
|
|
operators. External clients obtain reports only at finalization, not through debug text.
|
|
|
|
## Final public object
|
|
|
|
Every RunResult adds `operation_result`, null for an undeclared operation. Otherwise:
|
|
|
|
```json
|
|
{
|
|
"protocol": "aim_operation_result_v1",
|
|
"schema": "host_capabilities_v1",
|
|
"scope": "per_host",
|
|
"required": true,
|
|
"complete": true,
|
|
"check_mode": false,
|
|
"hosts": {
|
|
"host01.example": {
|
|
"schema": "host_capabilities_v1",
|
|
"status": "available",
|
|
"data": {"is_dc": false, "is_dhcp_server": false, "is_hyperv_host": false,
|
|
"has_veeam_vbr": false, "has_veeam_vbo": false, "has_veeam_em": false,
|
|
"is_unifi_controller": false, "is_unifi_os_server": false},
|
|
"error": null
|
|
}
|
|
},
|
|
"global": null
|
|
}
|
|
```
|
|
|
|
For global scope, `hosts` is empty and `global` holds the same entry shape. Schema and
|
|
scope come from the catalog, not client input. Host keys come from reviewed targets.
|
|
The existing final result event contains the same object as the final response. There
|
|
is no separate live operation-data event; read final authoritative results in either
|
|
summary or detail mode.
|
|
|
|
Entry status:
|
|
|
|
| Status | Meaning |
|
|
|---|---|
|
|
| available | Complete final data received, schema/limit/safety validation passed |
|
|
| missing | Final accounting arrived but this report slot was not published |
|
|
| withheld | Sensitive publisher or known supplied secret matched output |
|
|
| invalid | Schema, provenance or size validation rejected the publication |
|
|
| not_started | Execution did not start; there is no operation data |
|
|
| indeterminate | Work may have started but no complete final report stream exists |
|
|
|
|
Unavailable entries have `data: null`; never replace them with `{}`/zero/false in a UI.
|
|
`complete` describes report availability, not execution success or payload-specific
|
|
completeness. A patch report has its own `data.complete` for update evidence. A failed
|
|
service-start operation can have an available, complete report listing failed services.
|
|
|
|
If native exit is nonzero, preserve the native/Core failure and retain any valid reports
|
|
from the final stats. If native execution exits 0 but required output is missing,
|
|
withheld, invalid or incomplete, Core returns failed/result_validation with native
|
|
exit_code 0. Native target outcomes still describe their Ansible stats. This is an
|
|
intentional additional report-contract check, not a redefinition of native stats.
|
|
Optional missing reports do not fail an otherwise successful run; invalid submitted
|
|
reports do. Never replay automatically after result-validation failure.
|
|
|
|
## Generic schema subset and safety
|
|
|
|
The schema validator is standard-library-only, shared by callback and Core. Supported
|
|
JSON Schema keywords: `type` (including nullable type arrays), `properties`, `required`,
|
|
boolean `additionalProperties`, `items`, `enum`, `maxItems`, `maxLength`, `minimum`,
|
|
`maximum`, and documentation `description`. `$ref`, arbitrary validators and executable
|
|
schema extensions are not supported. Objects are closed unless explicitly declared open.
|
|
The only shipped open subtree is already parsed/redacted Checkmk `sections`.
|
|
|
|
Limits: maximum depth 20, 200,000 nodes per value, 20,000 items per collection, 8,192
|
|
characters per string, 1 MiB encoded envelope per host/global slot and 16 MiB per run.
|
|
Catalogs can lower the per-host byte cap. Data is chunked privately into <=2,048-byte
|
|
pieces; malformed sequences, duplicate publications, torn frames and oversize data are
|
|
rejected, never silently truncated. NaN/Infinity, binary objects, cycles and unknown
|
|
closed-object properties are rejected. Schemas and output labels must be static public
|
|
controller source.
|
|
|
|
Obvious secret field names are disallowed in schemas; redacted dynamic config keys
|
|
can only hold the exact `[REDACTED]` marker. Known supplied credentials matching public
|
|
strings are withheld again at the Core boundary. These measures are defense in depth,
|
|
not a general secret scanner. Config/other strings literally containing an undisclosed
|
|
secret remain an author/operator responsibility. Never publish environments, invocation,
|
|
exception objects, arbitrary debug values, commands or private keys.
|
|
|
|
`checkmk_user_config_v1` preserves parsed sections but redacts recognized password/token/
|
|
passphrase/credential/community fields, command/argument/environment bodies, MRPE
|
|
commands, URL credentials and obvious inline secrets. `redacted_paths` explains what
|
|
was withheld. Comments/formatting are not YAML data. The operation may read only basename
|
|
`check_mk.user.yml`, <=512 KiB, from a trusted configured path. It never modifies the file.
|
|
Unknown settings remain visible unless filtered; do not advertise this as guaranteed
|
|
secret-free content. No general arbitrary-file reader is added.
|
|
|
|
## Current schemas and exact semantic boundaries
|
|
|
|
Normative field shapes: `playbooks/schemas/<schema>.yml` (also in catalog metadata).
|
|
|
|
| Schema | Data meaning |
|
|
|---|---|
|
|
| host_capabilities_v1 | Eight booleans, source detection facts; no inventory membership edits |
|
|
| filesystem_usage_v1 | Selected operational mounts/volumes; byte quantities and nullable observations; Windows uses attached storage volumes from community.windows.win_disk_facts and excludes mapped/network drives |
|
|
| event_log_export_v1 | Channel names, age and target file paths; no log contents or controller download; check mode has no exported paths |
|
|
| service_start_summary_v1 | Before stopped services, policy eligibility, attempts, exclusions, actual after observations and fixed per-attempt failures |
|
|
| patch_summary_v1 | Per-package net version-set changes on Debian/RedHat, or Windows update IDs/titles/KBs with installed flags; never arbitrary manager dictionaries |
|
|
| managed_cleanup_preview_v1 | Managed filename candidates and observed deletions; no unknown-file purge |
|
|
| checkmk_user_config_v1 | Parsed/redacted sections and file metadata, not raw YAML/debug output |
|
|
| checkmk_agent_state_v1 | Installed status/version from registry/package query, actual service states, package/config/check changes |
|
|
| checkmk_agent_config_v1 | Named section/config/check change actions; no raw before/after values or file diffs |
|
|
|
|
### Services
|
|
|
|
`initially_stopped` is the observation before action. `eligible` applies auto/delayed-start
|
|
and configured include/exclude policy. `attempted` excludes skipped and check-mode actions.
|
|
`newly_running` is the observed stopped-to-started transition, including a concurrent
|
|
external start; `started_count` counts only attempted services now running. It is not
|
|
proof that the start return alone kept a service alive. `still_stopped` includes excluded
|
|
services; `unobserved` means missing after observations. Non-running attempted services
|
|
appear in `failed_to_start` with name, fixed reason/message and nullable numeric native
|
|
code. Failed reasons include dependency_failed, permission_denied, service_disabled,
|
|
start_timeout, service_not_found, logon_failed, start_failed, not_running_after_start and
|
|
state_unavailable. Localized unrecognized errors use generic start_failed; no raw text.
|
|
An unavailable post-query fails the run instead of inventing a successful observation.
|
|
|
|
### Packages and Checkmk
|
|
|
|
Linux queries native package databases before/after the update/reboot policy, preserving
|
|
architecture and parallel installed version sets. `updates` lists updated/installed/
|
|
removed rows and separate counters. Old/new arrays are empty where the package did not
|
|
exist. Snapshots cannot identify an intermediate reinstall with unchanged final versions
|
|
or attribute concurrent external package changes; do not run another package manager
|
|
concurrently. Check mode reports no completed changes, not predicted upgrade versions.
|
|
No silently shortened list: exceeding limits becomes explicit validation failure.
|
|
|
|
Windows reports native updates marked installed, pending/failed records and numeric
|
|
failure codes. A mismatch between native installed count and detailed records marks
|
|
`data.complete: false`, rather than fabricating records. Supported scope remains Windows
|
|
Update under configured categories, not every third-party package installer. Reboot
|
|
facts retain their observed meaning; check mode never claims a performed reboot.
|
|
|
|
Checkmk versions can be null (unavailable/ambiguous); do not fall back to the filename.
|
|
Config changes are section/file-level, not field-level raw diffs. `deployed_checks` means
|
|
selected file-copy tasks completed; `changed_checks` identifies changed copy results.
|
|
Check-mode changes are predictions; mode is always explicit. Unknown files are neither
|
|
counted nor scanned: policy is `untouched_not_enumerated`, not an invented preserved count.
|
|
|
|
## Errors and consumer behavior
|
|
|
|
Report errors: operation_result_missing, operation_result_invalid,
|
|
operation_result_withheld, operation_result_limit, operation_result_incomplete.
|
|
Messages contain no payload/exception content. Preserve final Core/native status;
|
|
show report issues separately; no credential retry or task replay is implied.
|
|
|
|
Treat this candidate as an implementation awaiting native 2.19.11/managed-host acceptance.
|
|
Upstream interfaces consulted: Ansible 2.19.11 `plugins/action/set_stats.py`,
|
|
`executor/stats.py`, and callback result contracts; ansible.windows >=3.8.0,<4.0.0 win_updates and
|
|
win_service_info return definitions. Consultation is not native testing.
|
|
|
|
|
|
## Patch wave and reboot semantics (3.3.0rc8)
|
|
|
|
`patch_summary_v1` reports package/update evidence independently from reboot policy. Linux
|
|
retains the rc2 package-manager transaction model. Windows delegates each selected patch
|
|
wave to one native `ansible.windows.win_updates` `state: installed` invocation with
|
|
`reboot: false`. The collection and Windows Update Agent own update ordering/coordination
|
|
inside that wave; AIM does not schedule updates individually.
|
|
|
|
The returned update dictionary is normalized into reviewed installed/failed records. A
|
|
module-level failure with no per-update records becomes a bounded generic failed-wave record;
|
|
raw failure text, exceptions and arbitrary module dictionaries are not published.
|
|
|
|
A newly required reboot with `os_patching_reboot: false` does not invalidate successful
|
|
installation: `reboot_required`/`reboot_required_after` are true and `reboot_deferred` is
|
|
true. The run stops. On a later run, a detected pre-existing reboot with automatic reboot
|
|
still disabled publishes `evidence: preflight_reboot_state`, `complete: false`,
|
|
`blocked_reason: preexisting_reboot_required`, then fails before new patch work starts.
|
|
|
|
When automatic reboot is enabled, AIM waits for the native update wave to return before
|
|
performing its reviewed message/delay/reboot. By default
|
|
`os_patching_rescan_after_reboot: false`; any AIM-performed reboot ends the run with
|
|
`continuation_required: true` and `remaining_updates_known: false`. This includes a reboot
|
|
that was already required before patching began. A new operator-approved run owns discovery
|
|
of the next patch state.
|
|
|
|
If the operator explicitly sets post-reboot continuation true, AIM may invoke another native
|
|
Windows update wave after reboot. Continuation is bounded to 12 wave invocations per run. A
|
|
failed wave is never automatically replayed or used as authority to continue.
|
|
|
|
When a Windows wave completes without reboot, AIM performs one final read-only search. This
|
|
does not install newly applicable updates. `remaining_updates_known` is then true and
|
|
`pending` contains that final observation; `continuation_required` indicates whether another
|
|
operator-approved run has applicable work.
|
|
|
|
`patch_cycles` counts native Windows install-wave invocations entered in the run and
|
|
`rescan_after_reboot` echoes the reviewed option. Those fields are Windows-specific and are
|
|
optional in the closed schema so existing Linux result shape remains stable.
|
|
|
|
Windows `failed_updates` contains only bounded safe fields: update identity/title, normalized
|
|
unsigned HRESULT and hex form, a fixed reason and fixed message. Known reasons include
|
|
`operation_in_progress`, `install_not_allowed`, `not_applicable`,
|
|
`exclusive_install_conflict`, `self_update_in_progress`, `no_connection`, and `timeout`.
|
|
Raw Windows Update failure strings are not published. `0x80240016` maps to
|
|
`install_not_allowed`; it does not prove a reboot was pending because Windows also uses that
|
|
HRESULT while another installation is active. Only the independent preflight can report
|
|
`preexisting_reboot_required`.
|
|
|
|
`reboot_required_before` records the pre-run observation. `reboot_performed` is true when
|
|
AIM completed an approved reboot. `reboot_required_after` and the legacy
|
|
`reboot_required` field describe the final pending state represented by the report.
|
|
`reboot_delay_minutes` echoes the reviewed delay. The user-facing reboot message itself is
|
|
not copied into the structured result.
|
|
|
|
Windows, Debian and RedHat use platform-native reboot-required signals; RedHat preflight
|
|
is conservative when `needs-restarting` is not already present. Report fields never imply
|
|
that every possible vendor-specific reboot indicator was discovered.
|
|
|