Files
Ansible/scripts/docs/DETAILED_PROGRESS.md
2026-09-22 19:23:17 +02:00

16 KiB

AIM 3.3.0rc8: safe play/task/host progress

Service / wire / event version: 1.0, stable additive 1.x. Detail schema: play_task_host_v1. Canonical Ansible Core: 2.19.11. Qualification: locally tested with simulated native command/callback fixtures; real Ansible 2.19.11 and managed-host acceptance is still required for this feature.

1. Select the feature without breaking existing clients

Clients receive anonymous progress and aggregate counters by omitting progress_mode. Current Core accepts summary (the default) or detail in RunRequest. There is no -vvv, debug-output forwarding or raw-output option.

Inspect aimctl capabilities / AimService.capabilities() first:

{
  "execution_progress": {
    "modes": ["summary", "detail"],
    "default": "summary",
    "request_field": "progress_mode",
    "detail_schema": "play_task_host_v1",
    "detail_event_kinds": ["play_started", "play_skipped", "play_stopped", "task_started", "host_result", "task_retry", "task_async_poll", "host_recap"],
    "static_source_labels_only": true,
    "raw_output": false,
    "error_policy": "fixed_diagnostic_hints",
    "qualification": "controller_acceptance_required"
  }
}

This is the relevant capability fragment, not the entire response. Servers without the advertised detail capability can reject unknown request fields: do not send progress_mode to them. Absence of the capability means use the legacy default or explain that detailed progress needs a newer core. Administrative enablement and same-UID runtime readiness remain separate checks.

For a known inventory host, prepare a request with detail selected:

printf '%s\n' '{"api_version":"1.0","operation":"prepare","request":{"customer":"CUSTOMER","playbook":"debug_test_connection","hosts":["HOST"],"progress_mode":"detail"}}' | aimctl request

Substitute real catalog/customer/host identifiers. Input to aimctl request is one JSON object on one line, not pretty-printed multiline JSON. Preparation collects no password and contacts no managed host. Use the same normalized request (including the mode) and its revision for execution. Changing the mode changes the reviewed request and requires preparation/review again. An ordinary request never contains credentials; use the private provider/FD documented in ADDON_API.md.

Detailed mode still emits existing stage, progress, stats, and result events. A detail renderer should ignore anonymous progress events to avoid duplicate task lines. An existing summary client can ignore unfamiliar event kinds. Product version is 3.3.0rc8; API, wire and event version strings stay 1.0.

2. Public event envelope and fields

Every event has event_version, run_id, increasing sequence, UTC timestamp, and kind. On the wire it is wrapped as {"type":"event","event":{...}}. IDs (play_id, task_id) are opaque strings scoped to a run. Correlate with IDs, not names or the last task that happened to arrive. Results can interleave across hosts under a free strategy. Serial batches can enter the same play again and get new IDs. Repeated task labels are not unique identities.

Kind Additional fields Meaning
play_started play_id, label, label_redacted A play occurrence began; label is unexpanded source text or a fixed placeholder
play_skipped play_id, reason: "no_hosts_matched" Native no-matching-host callback; not an authentication failure
play_stopped play_id, reason: "no_hosts_remaining" Native no-hosts-remaining callback; not equivalent to a harmless no-match skip
task_started play_id, task_id, label, label_redacted, handler A task/handler started; no arguments, module name, role path, source path or templated name
host_result play_id, task_id, host, host_redacted, status, changed, ignored, details_redacted, error One aggregate host/task outcome
task_retry play_id, task_id, host, host_redacted, attempt, details_redacted Native task retry, not AIM replay of the playbook
task_async_poll Same fields as task_retry Async poll notification; no job ID or result payload
host_recap host, host_redacted, counts Final counts for one logical inventory host

host_result.status is ok, changed, skipped, failed, or unreachable. changed is a boolean and can also be true on a failed task. ignored reflects native ignore-error/ignore-unreachable handling; a failed event is not by itself proof that the whole run failed. Rescue and ignored counts remain in final stats. attempt is an integer or null; it is always null when details are withheld.

host is a reviewed logical inventory name, never a resolved connection address, delegation target, loop label or variable. Out-of-scope/unsafe names are represented as null with host_redacted: true. This can occur for trusted delegated/dynamic inventory work: host limits are an operational selection, not a security sandbox. A host name containing a supplied credential is also withheld.

Every counts mapping has nonnegative integers for ok, changed, failures, unreachable, skipped, rescued, ignored. Host recaps are followed by the existing aggregate stats and final run result. Do not count progress lines to calculate a recap, assume one task header per host, or infer success from a zero process exit without the authoritative final response/result. A failure followed by rescue can still result in a successful native run.

There is no invented play_completed event: use the actual following play/final stats boundary. Loops produce aggregate host/task outcomes, not item values/events. Async poll and retry are optional native notifications; fire-and-forget async work is not thereby certified complete. The console's exact visual layout is not the API.

Example host outcome (IDs/timestamp illustrative):

{
  "event_version": "1.0",
  "run_id": "example-run",
  "sequence": 12,
  "timestamp": "2026-09-19T12:00:00+00:00",
  "kind": "host_result",
  "play_id": "p2",
  "task_id": "t1",
  "host": "host01.example.test",
  "host_redacted": false,
  "status": "unreachable",
  "changed": false,
  "ignored": false,
  "details_redacted": false,
  "error": {
    "code": "connection_refused",
    "message": "The connection was refused. Check the target listener, port and firewall.",
    "classification": "diagnostic_hint"
  }
}

3. Failure information without raw error text

A host failure carries error: {code, message, classification}. Success/skip events carry error: null. Messages are fixed strings owned by core. They never contain the original result's msg, module arguments, exception, stdout or stderr.

Code Interpretation
connection_refused Native unreachable message matched a connection-refusal signature; check listener, port and firewall
connection_timeout Native unreachable message matched a connection timeout
name_resolution_failed Native unreachable message matched a DNS/name-resolution error
tls_verification_failed Native unreachable message matched certificate verification failure
authentication_failed Native unreachable message matched a recognized authentication rejection
permission_denied Failed task message matched an access/permission denial
host_unreachable No narrower supported hint; transport/authentication details remain withheld
task_failed No narrower supported hint; arbitrary module error text remains withheld
details_withheld Sensitive/no_log result; no diagnostic classification is exposed

Hints are derived from bounded native message signatures, not a definitive root cause or proof that a supplied password was used. Localized/unrecognized messages can produce generic errors. Never drive automatic credential retries, disable TLS validation, open firewall rules or replay operations from these hints. The existing RunResult error/remote-work flag remains authoritative for lifecycle decisions. Preflight errors still use the existing fixed service error codes; this is not a raw syntax-error/traceback channel. More detail may require an authorized operator's trusted terminal diagnostics, handled as potentially sensitive data.

4. Safety contract and unavoidable trust boundary

Core reads the original parsed play/task name field, not Ansible's templated get_name() or rendered task fields. Missing, templated, overlong, control-bearing, URL-bearing or credential-assignment-like labels are replaced by fixed labels. Known supplied credential values matching a label are also withheld at the public boundary. Labels have at most 200 characters / 800 UTF-8 bytes; public host names at most 255 characters. Longer valid inventory identifiers can execute, but their name is withheld in the detail stream.

Known true or potentially true no_log on a task, block, role or play suppresses its label. For dynamic no_log expressions core does not attempt to render the expression. Results marked no_log, censored, or containing hidden loop results retain only status/changed/ignored, allowed host and IDs. Their failures use details_withheld. Retry attempts are hidden for these results.

Static source labels and inventory names must themselves be non-secret. No filter can identify every secret literally hard-coded in a name. A runtime-only sensitivity flag also cannot retroactively retract a static header already sent. This is a trusted controller/source contract, not a general secret-scanning engine or malicious-playbook sandbox. Add-on authors must never put credentials or variable values in labels. Dynamic names deliberately lose their expansion.

Nothing here authorizes raw debug values, Checkmk configuration contents, registered variables, invocation, environment/command strings, loop items, exceptions, custom stats, module stdout/stderr, or -v/-vv/-vvv/-vvvv output. raw_task_output remains unsupported. Known password masking is defense in depth, not permission to pass arbitrary output through a redactor.

Core validates the private callback schema, IDs, sequences and counts before building public events. Detail collection is bounded to 200,000 private frames (including legacy counters/handshakes), 50,000 play occurrences and 50,000 task IDs. These are safety bounds, not an estimated progress denominator. Malformed, truncated, out-of-order or incomplete streams cannot produce a successful run. Typical core errors are invalid_event_stream, event_limit, event_bridge_unavailable, and event_bridge_incomplete. A native failure may end before a complete recap: retain received events and show failure/unknown, never fabricate the missing recap.

The event sink must remain responsive or enqueue into a bounded queue; it must not block on slow browsers. Core still uses the established process deadlines and cancellation. No output detail mode changes account permissions, credential precedence, execution enablement, or automatic retry policy.

5. Example plain-text renderer for an independent client

This is an illustrative client function, not a new core CLI command. Feed it validated public event objects from your client after version/capability checks. Use text nodes/HTML escaping in a browser, not innerHTML; escape rich terminal markup as well. Treat labels as text, not as commands or markup. Persist only the minimal events your deployment policy needs and restrict job-log visibility.

class ProgressText:
    def __init__(self):
        self.tasks = {}
        self.active_task = None
        self.recap_started = False

    def __call__(self, event):
        kind = event.get('kind')
        if kind == 'play_started':
            print('\nPLAY [' + event['label'] + ']', flush=True)
            self.active_task = None
        elif kind == 'play_skipped':
            print('skipping: no hosts matched', flush=True)
        elif kind == 'play_stopped':
            print('stopped: no hosts remaining', flush=True)
        elif kind == 'task_started':
            self.tasks[event['task_id']] = event['label']
            self.active_task = event['task_id']
            prefix = 'HANDLER' if event['handler'] else 'TASK'
            print('\n' + prefix + ' [' + event['label'] + ']', flush=True)
        elif kind == 'host_result':
            task_id = event['task_id']
            # Results can interleave: repeat the correct header, not the last name.
            if self.active_task != task_id:
                print('\nTASK [' + self.tasks.get(task_id, 'Task') + ']', flush=True)
                self.active_task = task_id
            host = event['host'] or 'host withheld'
            suffix = ' (ignored)' if event['ignored'] else ''
            error = event.get('error')
            if error:
                suffix += ' => ' + error['message']
            print(event['status'] + ': [' + host + ']' + suffix, flush=True)
        elif kind in ('task_retry', 'task_async_poll'):
            print(kind + ': [' + (event['host'] or 'host withheld') + ']', flush=True)
        elif kind == 'host_recap':
            if not self.recap_started:
                print('\nPLAY RECAP', flush=True)
                self.recap_started = True
            counts = ' '.join(k + '=' + str(v) for k, v in event['counts'].items())
            print((event['host'] or 'host withheld') + ' : ' + counts, flush=True)
        # Ignore legacy anonymous progress to avoid duplicate lines in detail mode.
        # The caller handles stage/stats/result and the authoritative final response.

A GUI timeline should key results by (run_id, play_id, task_id, host) where the host is visible. It must not merge all withheld-host records as one known machine. Keep the original event sequence for audit; a display regrouping is presentation only. The built-in aim terminal still uses its existing native Ansible output, not this example renderer. No add-on source was changed for this feature.

6. Native contracts consulted (not proof of native execution)

Implementation targets Ansible Core 2.19.11 callback entry points, original source mappings and CallbackTaskResult public properties. Any dependence on these internals stays in the core-owned adapter, not in an add-on.

Follow SANITY.md for exact-runtime/controller acceptance before treating the new timeline as qualified in an installed service or add-on deployment.

Final target outcomes

Do not reconstruct final requested-target success from this live detail stream. RunResult.target_summary and RunResult.targets now provide Core-owned final accounting under target_outcome_summary_v1 in both summary and detail modes. The detail stream remains for live presentation and diagnostics; the final target summary is authoritative for per-request host outcomes. See TARGET_OUTCOMES.md.

Purposeful data is separate

The 3.3.0rc8 operation_result channel does not loosen this host/task event contract. Only explicitly catalogued, typed and validated set_stats data is returned in the final result; see OPERATION_RESULTS.md. Arbitrary debug output and configuration bodies remain excluded from progress events. The Checkmk reader now publishes parsed redacted sections through its own declared report, not raw debug text.