aim-web2.1.0rc9
This commit is contained in:
@@ -0,0 +1,284 @@
|
||||
# AIM 3.3.0rc8: safe play/task/host progress
|
||||
|
||||
**Service / wire / event version:** 1.0, stable additive 1.x.
|
||||
**Detail schema:** `play_task_host_v1`. **Canonical Ansible Core:** 2.19.11.
|
||||
**Qualification:** locally tested with simulated native command/callback fixtures;
|
||||
real Ansible 2.19.11 and managed-host acceptance is still required for this feature.
|
||||
|
||||
## 1. Select the feature without breaking existing clients
|
||||
|
||||
Clients receive anonymous progress and aggregate counters by omitting
|
||||
`progress_mode`. Current Core accepts `summary` (the default) or `detail` in
|
||||
`RunRequest`. There is no `-vvv`, debug-output forwarding or raw-output option.
|
||||
|
||||
Inspect `aimctl capabilities` / `AimService.capabilities()` first:
|
||||
|
||||
```json
|
||||
{
|
||||
"execution_progress": {
|
||||
"modes": ["summary", "detail"],
|
||||
"default": "summary",
|
||||
"request_field": "progress_mode",
|
||||
"detail_schema": "play_task_host_v1",
|
||||
"detail_event_kinds": ["play_started", "play_skipped", "play_stopped", "task_started", "host_result", "task_retry", "task_async_poll", "host_recap"],
|
||||
"static_source_labels_only": true,
|
||||
"raw_output": false,
|
||||
"error_policy": "fixed_diagnostic_hints",
|
||||
"qualification": "controller_acceptance_required"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
This is the relevant capability fragment, not the entire response. Servers without
|
||||
the advertised detail capability can reject unknown request fields: do not send
|
||||
`progress_mode` to them. Absence of the capability means use the legacy default or
|
||||
explain that detailed progress needs a newer core. Administrative enablement and
|
||||
same-UID runtime readiness remain separate checks.
|
||||
|
||||
For a known inventory host, prepare a request with detail selected:
|
||||
|
||||
```bash
|
||||
printf '%s\n' '{"api_version":"1.0","operation":"prepare","request":{"customer":"CUSTOMER","playbook":"debug_test_connection","hosts":["HOST"],"progress_mode":"detail"}}' | aimctl request
|
||||
```
|
||||
|
||||
Substitute real catalog/customer/host identifiers. Input to `aimctl request` is
|
||||
**one JSON object on one line**, not pretty-printed multiline JSON. Preparation
|
||||
collects no password and contacts no managed host. Use the same normalized request
|
||||
(including the mode) and its revision for execution. Changing the mode changes the
|
||||
reviewed request and requires preparation/review again. An ordinary request never
|
||||
contains credentials; use the private provider/FD documented in `ADDON_API.md`.
|
||||
|
||||
Detailed mode still emits existing `stage`, `progress`, `stats`, and `result` events.
|
||||
A detail renderer should ignore anonymous `progress` events to avoid duplicate task
|
||||
lines. An existing summary client can ignore unfamiliar event kinds. Product version
|
||||
is 3.3.0rc8; API, wire and event version strings stay 1.0.
|
||||
|
||||
## 2. Public event envelope and fields
|
||||
|
||||
Every event has `event_version`, `run_id`, increasing `sequence`, UTC `timestamp`,
|
||||
and `kind`. On the wire it is wrapped as `{"type":"event","event":{...}}`.
|
||||
IDs (`play_id`, `task_id`) are opaque strings scoped to a run. Correlate with IDs,
|
||||
not names or the last task that happened to arrive. Results can interleave across
|
||||
hosts under a free strategy. Serial batches can enter the same play again and get
|
||||
new IDs. Repeated task labels are not unique identities.
|
||||
|
||||
| Kind | Additional fields | Meaning |
|
||||
|---|---|---|
|
||||
| `play_started` | `play_id`, `label`, `label_redacted` | A play occurrence began; label is unexpanded source text or a fixed placeholder |
|
||||
| `play_skipped` | `play_id`, `reason: "no_hosts_matched"` | Native no-matching-host callback; not an authentication failure |
|
||||
| `play_stopped` | `play_id`, `reason: "no_hosts_remaining"` | Native no-hosts-remaining callback; not equivalent to a harmless no-match skip |
|
||||
| `task_started` | `play_id`, `task_id`, `label`, `label_redacted`, `handler` | A task/handler started; no arguments, module name, role path, source path or templated name |
|
||||
| `host_result` | `play_id`, `task_id`, `host`, `host_redacted`, `status`, `changed`, `ignored`, `details_redacted`, `error` | One aggregate host/task outcome |
|
||||
| `task_retry` | `play_id`, `task_id`, `host`, `host_redacted`, `attempt`, `details_redacted` | Native task retry, not AIM replay of the playbook |
|
||||
| `task_async_poll` | Same fields as `task_retry` | Async poll notification; no job ID or result payload |
|
||||
| `host_recap` | `host`, `host_redacted`, `counts` | Final counts for one logical inventory host |
|
||||
|
||||
`host_result.status` is `ok`, `changed`, `skipped`, `failed`, or `unreachable`.
|
||||
`changed` is a boolean and can also be true on a failed task. `ignored` reflects
|
||||
native ignore-error/ignore-unreachable handling; a failed event is not by itself
|
||||
proof that the whole run failed. Rescue and ignored counts remain in final stats.
|
||||
`attempt` is an integer or null; it is always null when details are withheld.
|
||||
|
||||
`host` is a reviewed logical inventory name, never a resolved connection address,
|
||||
delegation target, loop label or variable. Out-of-scope/unsafe names are represented
|
||||
as null with `host_redacted: true`. This can occur for trusted delegated/dynamic
|
||||
inventory work: host limits are an operational selection, not a security sandbox.
|
||||
A host name containing a supplied credential is also withheld.
|
||||
|
||||
Every `counts` mapping has nonnegative integers for `ok`, `changed`, `failures`,
|
||||
`unreachable`, `skipped`, `rescued`, `ignored`. Host recaps are followed by the
|
||||
existing aggregate `stats` and final run result. Do not count progress lines to
|
||||
calculate a recap, assume one task header per host, or infer success from a zero
|
||||
process exit without the authoritative final response/result. A failure followed
|
||||
by rescue can still result in a successful native run.
|
||||
|
||||
There is no invented `play_completed` event: use the actual following play/final
|
||||
stats boundary. Loops produce aggregate host/task outcomes, not item values/events.
|
||||
Async poll and retry are optional native notifications; fire-and-forget async work
|
||||
is not thereby certified complete. The console's exact visual layout is not the API.
|
||||
|
||||
Example host outcome (IDs/timestamp illustrative):
|
||||
|
||||
```json
|
||||
{
|
||||
"event_version": "1.0",
|
||||
"run_id": "example-run",
|
||||
"sequence": 12,
|
||||
"timestamp": "2026-09-19T12:00:00+00:00",
|
||||
"kind": "host_result",
|
||||
"play_id": "p2",
|
||||
"task_id": "t1",
|
||||
"host": "host01.example.test",
|
||||
"host_redacted": false,
|
||||
"status": "unreachable",
|
||||
"changed": false,
|
||||
"ignored": false,
|
||||
"details_redacted": false,
|
||||
"error": {
|
||||
"code": "connection_refused",
|
||||
"message": "The connection was refused. Check the target listener, port and firewall.",
|
||||
"classification": "diagnostic_hint"
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
## 3. Failure information without raw error text
|
||||
|
||||
A host failure carries `error: {code, message, classification}`. Success/skip events
|
||||
carry `error: null`. Messages are fixed strings owned by core. They never contain
|
||||
the original result's `msg`, module arguments, exception, stdout or stderr.
|
||||
|
||||
| Code | Interpretation |
|
||||
|---|---|
|
||||
| `connection_refused` | Native unreachable message matched a connection-refusal signature; check listener, port and firewall |
|
||||
| `connection_timeout` | Native unreachable message matched a connection timeout |
|
||||
| `name_resolution_failed` | Native unreachable message matched a DNS/name-resolution error |
|
||||
| `tls_verification_failed` | Native unreachable message matched certificate verification failure |
|
||||
| `authentication_failed` | Native unreachable message matched a recognized authentication rejection |
|
||||
| `permission_denied` | Failed task message matched an access/permission denial |
|
||||
| `host_unreachable` | No narrower supported hint; transport/authentication details remain withheld |
|
||||
| `task_failed` | No narrower supported hint; arbitrary module error text remains withheld |
|
||||
| `details_withheld` | Sensitive/no_log result; no diagnostic classification is exposed |
|
||||
|
||||
Hints are derived from bounded native message signatures, **not a definitive root
|
||||
cause or proof that a supplied password was used**. Localized/unrecognized messages
|
||||
can produce generic errors. Never drive automatic credential retries, disable TLS
|
||||
validation, open firewall rules or replay operations from these hints. The existing
|
||||
RunResult error/remote-work flag remains authoritative for lifecycle decisions.
|
||||
Preflight errors still use the existing fixed service error codes; this is not a
|
||||
raw syntax-error/traceback channel. More detail may require an authorized operator's
|
||||
trusted terminal diagnostics, handled as potentially sensitive data.
|
||||
|
||||
## 4. Safety contract and unavoidable trust boundary
|
||||
|
||||
Core reads the original parsed play/task `name` field, not Ansible's templated
|
||||
`get_name()` or rendered task fields. Missing, templated, overlong, control-bearing,
|
||||
URL-bearing or credential-assignment-like labels are replaced by fixed labels.
|
||||
Known supplied credential values matching a label are also withheld at the public
|
||||
boundary. Labels have at most 200 characters / 800 UTF-8 bytes; public host names
|
||||
at most 255 characters. Longer valid inventory identifiers can execute, but their
|
||||
name is withheld in the detail stream.
|
||||
|
||||
Known true or potentially true `no_log` on a task, block, role or play suppresses
|
||||
its label. For dynamic no_log expressions core does not attempt to render the
|
||||
expression. Results marked `no_log`, censored, or containing hidden loop results
|
||||
retain only status/changed/ignored, allowed host and IDs. Their failures use
|
||||
`details_withheld`. Retry attempts are hidden for these results.
|
||||
|
||||
**Static source labels and inventory names must themselves be non-secret.** No
|
||||
filter can identify every secret literally hard-coded in a name. A runtime-only
|
||||
sensitivity flag also cannot retroactively retract a static header already sent.
|
||||
This is a trusted controller/source contract, not a general secret-scanning engine
|
||||
or malicious-playbook sandbox. Add-on authors must never put credentials or
|
||||
variable values in labels. Dynamic names deliberately lose their expansion.
|
||||
|
||||
Nothing here authorizes raw `debug` values, Checkmk configuration contents,
|
||||
registered variables, `invocation`, environment/command strings, loop items,
|
||||
exceptions, custom stats, module stdout/stderr, or `-v/-vv/-vvv/-vvvv` output.
|
||||
`raw_task_output` remains unsupported. Known password masking is defense in depth,
|
||||
not permission to pass arbitrary output through a redactor.
|
||||
|
||||
Core validates the private callback schema, IDs, sequences and counts before
|
||||
building public events. Detail collection is bounded to 200,000 private frames
|
||||
(including legacy counters/handshakes), 50,000 play occurrences and 50,000 task IDs.
|
||||
These are safety bounds, not an estimated progress denominator. Malformed, truncated,
|
||||
out-of-order or incomplete streams cannot produce a successful run. Typical core
|
||||
errors are `invalid_event_stream`, `event_limit`, `event_bridge_unavailable`, and
|
||||
`event_bridge_incomplete`. A native failure may end before a complete recap: retain
|
||||
received events and show failure/unknown, never fabricate the missing recap.
|
||||
|
||||
The event sink must remain responsive or enqueue into a bounded queue; it must not
|
||||
block on slow browsers. Core still uses the established process deadlines and
|
||||
cancellation. No output detail mode changes account permissions, credential
|
||||
precedence, execution enablement, or automatic retry policy.
|
||||
|
||||
## 5. Example plain-text renderer for an independent client
|
||||
|
||||
This is an illustrative client function, not a new core CLI command. Feed it
|
||||
validated public event objects from your client after version/capability checks.
|
||||
Use text nodes/HTML escaping in a browser, not `innerHTML`; escape rich terminal
|
||||
markup as well. Treat labels as text, not as commands or markup. Persist only the
|
||||
minimal events your deployment policy needs and restrict job-log visibility.
|
||||
|
||||
```python
|
||||
class ProgressText:
|
||||
def __init__(self):
|
||||
self.tasks = {}
|
||||
self.active_task = None
|
||||
self.recap_started = False
|
||||
|
||||
def __call__(self, event):
|
||||
kind = event.get('kind')
|
||||
if kind == 'play_started':
|
||||
print('\nPLAY [' + event['label'] + ']', flush=True)
|
||||
self.active_task = None
|
||||
elif kind == 'play_skipped':
|
||||
print('skipping: no hosts matched', flush=True)
|
||||
elif kind == 'play_stopped':
|
||||
print('stopped: no hosts remaining', flush=True)
|
||||
elif kind == 'task_started':
|
||||
self.tasks[event['task_id']] = event['label']
|
||||
self.active_task = event['task_id']
|
||||
prefix = 'HANDLER' if event['handler'] else 'TASK'
|
||||
print('\n' + prefix + ' [' + event['label'] + ']', flush=True)
|
||||
elif kind == 'host_result':
|
||||
task_id = event['task_id']
|
||||
# Results can interleave: repeat the correct header, not the last name.
|
||||
if self.active_task != task_id:
|
||||
print('\nTASK [' + self.tasks.get(task_id, 'Task') + ']', flush=True)
|
||||
self.active_task = task_id
|
||||
host = event['host'] or 'host withheld'
|
||||
suffix = ' (ignored)' if event['ignored'] else ''
|
||||
error = event.get('error')
|
||||
if error:
|
||||
suffix += ' => ' + error['message']
|
||||
print(event['status'] + ': [' + host + ']' + suffix, flush=True)
|
||||
elif kind in ('task_retry', 'task_async_poll'):
|
||||
print(kind + ': [' + (event['host'] or 'host withheld') + ']', flush=True)
|
||||
elif kind == 'host_recap':
|
||||
if not self.recap_started:
|
||||
print('\nPLAY RECAP', flush=True)
|
||||
self.recap_started = True
|
||||
counts = ' '.join(k + '=' + str(v) for k, v in event['counts'].items())
|
||||
print((event['host'] or 'host withheld') + ' : ' + counts, flush=True)
|
||||
# Ignore legacy anonymous progress to avoid duplicate lines in detail mode.
|
||||
# The caller handles stage/stats/result and the authoritative final response.
|
||||
```
|
||||
|
||||
A GUI timeline should key results by `(run_id, play_id, task_id, host)` where the
|
||||
host is visible. It must not merge all withheld-host records as one known machine.
|
||||
Keep the original event sequence for audit; a display regrouping is presentation
|
||||
only. The built-in `aim` terminal still uses its existing native Ansible output,
|
||||
not this example renderer. No add-on source was changed for this feature.
|
||||
|
||||
## 6. Native contracts consulted (not proof of native execution)
|
||||
|
||||
Implementation targets Ansible Core 2.19.11 callback entry points, original source
|
||||
mappings and CallbackTaskResult public properties. Any dependence on these internals
|
||||
stays in the core-owned adapter, not in an add-on.
|
||||
|
||||
- https://docs.ansible.com/projects/ansible-core/2.19/plugins/callback.html
|
||||
- https://docs.ansible.com/projects/ansible-core/2.19/reference_appendices/logging.html
|
||||
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/plugins/callback/default.py
|
||||
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/executor/task_result.py
|
||||
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/executor/playbook_executor.py
|
||||
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/playbook/base.py
|
||||
|
||||
Follow `SANITY.md` for exact-runtime/controller acceptance before treating the
|
||||
new timeline as qualified in an installed service or add-on deployment.
|
||||
|
||||
## Final target outcomes
|
||||
|
||||
Do not reconstruct final requested-target success from this live detail stream.
|
||||
`RunResult.target_summary` and `RunResult.targets` now provide Core-owned final
|
||||
accounting under `target_outcome_summary_v1` in both summary and detail modes. The
|
||||
detail stream remains for live presentation and diagnostics; the final target
|
||||
summary is authoritative for per-request host outcomes. See `TARGET_OUTCOMES.md`.
|
||||
|
||||
## Purposeful data is separate
|
||||
|
||||
The 3.3.0rc8 `operation_result` channel does not loosen this host/task event contract.
|
||||
Only explicitly catalogued, typed and validated set_stats data is returned in the
|
||||
final result; see OPERATION_RESULTS.md. Arbitrary debug output and configuration
|
||||
bodies remain excluded from progress events. The Checkmk reader now publishes parsed
|
||||
redacted sections through its own declared report, not raw debug text.
|
||||
Reference in New Issue
Block a user