Files
Ansible/scripts/docs/DETAILED_PROGRESS.md
T
2026-09-22 19:23:17 +02:00

285 lines
16 KiB
Markdown

# AIM 3.3.0rc8: safe play/task/host progress
**Service / wire / event version:** 1.0, stable additive 1.x.
**Detail schema:** `play_task_host_v1`. **Canonical Ansible Core:** 2.19.11.
**Qualification:** locally tested with simulated native command/callback fixtures;
real Ansible 2.19.11 and managed-host acceptance is still required for this feature.
## 1. Select the feature without breaking existing clients
Clients receive anonymous progress and aggregate counters by omitting
`progress_mode`. Current Core accepts `summary` (the default) or `detail` in
`RunRequest`. There is no `-vvv`, debug-output forwarding or raw-output option.
Inspect `aimctl capabilities` / `AimService.capabilities()` first:
```json
{
"execution_progress": {
"modes": ["summary", "detail"],
"default": "summary",
"request_field": "progress_mode",
"detail_schema": "play_task_host_v1",
"detail_event_kinds": ["play_started", "play_skipped", "play_stopped", "task_started", "host_result", "task_retry", "task_async_poll", "host_recap"],
"static_source_labels_only": true,
"raw_output": false,
"error_policy": "fixed_diagnostic_hints",
"qualification": "controller_acceptance_required"
}
}
```
This is the relevant capability fragment, not the entire response. Servers without
the advertised detail capability can reject unknown request fields: do not send
`progress_mode` to them. Absence of the capability means use the legacy default or
explain that detailed progress needs a newer core. Administrative enablement and
same-UID runtime readiness remain separate checks.
For a known inventory host, prepare a request with detail selected:
```bash
printf '%s\n' '{"api_version":"1.0","operation":"prepare","request":{"customer":"CUSTOMER","playbook":"debug_test_connection","hosts":["HOST"],"progress_mode":"detail"}}' | aimctl request
```
Substitute real catalog/customer/host identifiers. Input to `aimctl request` is
**one JSON object on one line**, not pretty-printed multiline JSON. Preparation
collects no password and contacts no managed host. Use the same normalized request
(including the mode) and its revision for execution. Changing the mode changes the
reviewed request and requires preparation/review again. An ordinary request never
contains credentials; use the private provider/FD documented in `ADDON_API.md`.
Detailed mode still emits existing `stage`, `progress`, `stats`, and `result` events.
A detail renderer should ignore anonymous `progress` events to avoid duplicate task
lines. An existing summary client can ignore unfamiliar event kinds. Product version
is 3.3.0rc8; API, wire and event version strings stay 1.0.
## 2. Public event envelope and fields
Every event has `event_version`, `run_id`, increasing `sequence`, UTC `timestamp`,
and `kind`. On the wire it is wrapped as `{"type":"event","event":{...}}`.
IDs (`play_id`, `task_id`) are opaque strings scoped to a run. Correlate with IDs,
not names or the last task that happened to arrive. Results can interleave across
hosts under a free strategy. Serial batches can enter the same play again and get
new IDs. Repeated task labels are not unique identities.
| Kind | Additional fields | Meaning |
|---|---|---|
| `play_started` | `play_id`, `label`, `label_redacted` | A play occurrence began; label is unexpanded source text or a fixed placeholder |
| `play_skipped` | `play_id`, `reason: "no_hosts_matched"` | Native no-matching-host callback; not an authentication failure |
| `play_stopped` | `play_id`, `reason: "no_hosts_remaining"` | Native no-hosts-remaining callback; not equivalent to a harmless no-match skip |
| `task_started` | `play_id`, `task_id`, `label`, `label_redacted`, `handler` | A task/handler started; no arguments, module name, role path, source path or templated name |
| `host_result` | `play_id`, `task_id`, `host`, `host_redacted`, `status`, `changed`, `ignored`, `details_redacted`, `error` | One aggregate host/task outcome |
| `task_retry` | `play_id`, `task_id`, `host`, `host_redacted`, `attempt`, `details_redacted` | Native task retry, not AIM replay of the playbook |
| `task_async_poll` | Same fields as `task_retry` | Async poll notification; no job ID or result payload |
| `host_recap` | `host`, `host_redacted`, `counts` | Final counts for one logical inventory host |
`host_result.status` is `ok`, `changed`, `skipped`, `failed`, or `unreachable`.
`changed` is a boolean and can also be true on a failed task. `ignored` reflects
native ignore-error/ignore-unreachable handling; a failed event is not by itself
proof that the whole run failed. Rescue and ignored counts remain in final stats.
`attempt` is an integer or null; it is always null when details are withheld.
`host` is a reviewed logical inventory name, never a resolved connection address,
delegation target, loop label or variable. Out-of-scope/unsafe names are represented
as null with `host_redacted: true`. This can occur for trusted delegated/dynamic
inventory work: host limits are an operational selection, not a security sandbox.
A host name containing a supplied credential is also withheld.
Every `counts` mapping has nonnegative integers for `ok`, `changed`, `failures`,
`unreachable`, `skipped`, `rescued`, `ignored`. Host recaps are followed by the
existing aggregate `stats` and final run result. Do not count progress lines to
calculate a recap, assume one task header per host, or infer success from a zero
process exit without the authoritative final response/result. A failure followed
by rescue can still result in a successful native run.
There is no invented `play_completed` event: use the actual following play/final
stats boundary. Loops produce aggregate host/task outcomes, not item values/events.
Async poll and retry are optional native notifications; fire-and-forget async work
is not thereby certified complete. The console's exact visual layout is not the API.
Example host outcome (IDs/timestamp illustrative):
```json
{
"event_version": "1.0",
"run_id": "example-run",
"sequence": 12,
"timestamp": "2026-09-19T12:00:00+00:00",
"kind": "host_result",
"play_id": "p2",
"task_id": "t1",
"host": "host01.example.test",
"host_redacted": false,
"status": "unreachable",
"changed": false,
"ignored": false,
"details_redacted": false,
"error": {
"code": "connection_refused",
"message": "The connection was refused. Check the target listener, port and firewall.",
"classification": "diagnostic_hint"
}
}
```
## 3. Failure information without raw error text
A host failure carries `error: {code, message, classification}`. Success/skip events
carry `error: null`. Messages are fixed strings owned by core. They never contain
the original result's `msg`, module arguments, exception, stdout or stderr.
| Code | Interpretation |
|---|---|
| `connection_refused` | Native unreachable message matched a connection-refusal signature; check listener, port and firewall |
| `connection_timeout` | Native unreachable message matched a connection timeout |
| `name_resolution_failed` | Native unreachable message matched a DNS/name-resolution error |
| `tls_verification_failed` | Native unreachable message matched certificate verification failure |
| `authentication_failed` | Native unreachable message matched a recognized authentication rejection |
| `permission_denied` | Failed task message matched an access/permission denial |
| `host_unreachable` | No narrower supported hint; transport/authentication details remain withheld |
| `task_failed` | No narrower supported hint; arbitrary module error text remains withheld |
| `details_withheld` | Sensitive/no_log result; no diagnostic classification is exposed |
Hints are derived from bounded native message signatures, **not a definitive root
cause or proof that a supplied password was used**. Localized/unrecognized messages
can produce generic errors. Never drive automatic credential retries, disable TLS
validation, open firewall rules or replay operations from these hints. The existing
RunResult error/remote-work flag remains authoritative for lifecycle decisions.
Preflight errors still use the existing fixed service error codes; this is not a
raw syntax-error/traceback channel. More detail may require an authorized operator's
trusted terminal diagnostics, handled as potentially sensitive data.
## 4. Safety contract and unavoidable trust boundary
Core reads the original parsed play/task `name` field, not Ansible's templated
`get_name()` or rendered task fields. Missing, templated, overlong, control-bearing,
URL-bearing or credential-assignment-like labels are replaced by fixed labels.
Known supplied credential values matching a label are also withheld at the public
boundary. Labels have at most 200 characters / 800 UTF-8 bytes; public host names
at most 255 characters. Longer valid inventory identifiers can execute, but their
name is withheld in the detail stream.
Known true or potentially true `no_log` on a task, block, role or play suppresses
its label. For dynamic no_log expressions core does not attempt to render the
expression. Results marked `no_log`, censored, or containing hidden loop results
retain only status/changed/ignored, allowed host and IDs. Their failures use
`details_withheld`. Retry attempts are hidden for these results.
**Static source labels and inventory names must themselves be non-secret.** No
filter can identify every secret literally hard-coded in a name. A runtime-only
sensitivity flag also cannot retroactively retract a static header already sent.
This is a trusted controller/source contract, not a general secret-scanning engine
or malicious-playbook sandbox. Add-on authors must never put credentials or
variable values in labels. Dynamic names deliberately lose their expansion.
Nothing here authorizes raw `debug` values, Checkmk configuration contents,
registered variables, `invocation`, environment/command strings, loop items,
exceptions, custom stats, module stdout/stderr, or `-v/-vv/-vvv/-vvvv` output.
`raw_task_output` remains unsupported. Known password masking is defense in depth,
not permission to pass arbitrary output through a redactor.
Core validates the private callback schema, IDs, sequences and counts before
building public events. Detail collection is bounded to 200,000 private frames
(including legacy counters/handshakes), 50,000 play occurrences and 50,000 task IDs.
These are safety bounds, not an estimated progress denominator. Malformed, truncated,
out-of-order or incomplete streams cannot produce a successful run. Typical core
errors are `invalid_event_stream`, `event_limit`, `event_bridge_unavailable`, and
`event_bridge_incomplete`. A native failure may end before a complete recap: retain
received events and show failure/unknown, never fabricate the missing recap.
The event sink must remain responsive or enqueue into a bounded queue; it must not
block on slow browsers. Core still uses the established process deadlines and
cancellation. No output detail mode changes account permissions, credential
precedence, execution enablement, or automatic retry policy.
## 5. Example plain-text renderer for an independent client
This is an illustrative client function, not a new core CLI command. Feed it
validated public event objects from your client after version/capability checks.
Use text nodes/HTML escaping in a browser, not `innerHTML`; escape rich terminal
markup as well. Treat labels as text, not as commands or markup. Persist only the
minimal events your deployment policy needs and restrict job-log visibility.
```python
class ProgressText:
def __init__(self):
self.tasks = {}
self.active_task = None
self.recap_started = False
def __call__(self, event):
kind = event.get('kind')
if kind == 'play_started':
print('\nPLAY [' + event['label'] + ']', flush=True)
self.active_task = None
elif kind == 'play_skipped':
print('skipping: no hosts matched', flush=True)
elif kind == 'play_stopped':
print('stopped: no hosts remaining', flush=True)
elif kind == 'task_started':
self.tasks[event['task_id']] = event['label']
self.active_task = event['task_id']
prefix = 'HANDLER' if event['handler'] else 'TASK'
print('\n' + prefix + ' [' + event['label'] + ']', flush=True)
elif kind == 'host_result':
task_id = event['task_id']
# Results can interleave: repeat the correct header, not the last name.
if self.active_task != task_id:
print('\nTASK [' + self.tasks.get(task_id, 'Task') + ']', flush=True)
self.active_task = task_id
host = event['host'] or 'host withheld'
suffix = ' (ignored)' if event['ignored'] else ''
error = event.get('error')
if error:
suffix += ' => ' + error['message']
print(event['status'] + ': [' + host + ']' + suffix, flush=True)
elif kind in ('task_retry', 'task_async_poll'):
print(kind + ': [' + (event['host'] or 'host withheld') + ']', flush=True)
elif kind == 'host_recap':
if not self.recap_started:
print('\nPLAY RECAP', flush=True)
self.recap_started = True
counts = ' '.join(k + '=' + str(v) for k, v in event['counts'].items())
print((event['host'] or 'host withheld') + ' : ' + counts, flush=True)
# Ignore legacy anonymous progress to avoid duplicate lines in detail mode.
# The caller handles stage/stats/result and the authoritative final response.
```
A GUI timeline should key results by `(run_id, play_id, task_id, host)` where the
host is visible. It must not merge all withheld-host records as one known machine.
Keep the original event sequence for audit; a display regrouping is presentation
only. The built-in `aim` terminal still uses its existing native Ansible output,
not this example renderer. No add-on source was changed for this feature.
## 6. Native contracts consulted (not proof of native execution)
Implementation targets Ansible Core 2.19.11 callback entry points, original source
mappings and CallbackTaskResult public properties. Any dependence on these internals
stays in the core-owned adapter, not in an add-on.
- https://docs.ansible.com/projects/ansible-core/2.19/plugins/callback.html
- https://docs.ansible.com/projects/ansible-core/2.19/reference_appendices/logging.html
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/plugins/callback/default.py
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/executor/task_result.py
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/executor/playbook_executor.py
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/playbook/base.py
Follow `SANITY.md` for exact-runtime/controller acceptance before treating the
new timeline as qualified in an installed service or add-on deployment.
## Final target outcomes
Do not reconstruct final requested-target success from this live detail stream.
`RunResult.target_summary` and `RunResult.targets` now provide Core-owned final
accounting under `target_outcome_summary_v1` in both summary and detail modes. The
detail stream remains for live presentation and diagnostics; the final target
summary is authoritative for per-request host outcomes. See `TARGET_OUTCOMES.md`.
## Purposeful data is separate
The 3.3.0rc8 `operation_result` channel does not loosen this host/task event contract.
Only explicitly catalogued, typed and validated set_stats data is returned in the
final result; see OPERATION_RESULTS.md. Arbitrary debug output and configuration
bodies remain excluded from progress events. The Checkmk reader now publishes parsed
redacted sections through its own declared report, not raw debug text.