aim-web2.1.0rc9

This commit is contained in:
admin_rb
2026-09-22 19:23:17 +02:00
parent d095887d2e
commit 3dfc80b782
438 changed files with 31613 additions and 1510 deletions
+105
View File
@@ -0,0 +1,105 @@
# AIM add-on development contract
**Core 3.3.0rc8; service/wire/event 1.0 (stable additive 1.x); Ansible Core 2.19.11.**
AIM is the independently usable Source of Truth. Add-ons own presentation, sessions,
authorization at their application boundary, approvals, queues and persistence.
They never patch Core, scrape terminal output or duplicate inventory/execution rules.
## Integration sequence
Use `aim.services.v1` or one `aimctl request` process per JSON-lines request. Discover
capabilities, current customers/hosts/hierarchy and catalog. Prepare explicit hosts and
declared options, show the normalized request and result contract for review, check
readiness under the actual worker identity, acquire a worker slot, collect fresh
credentials, and execute using the prepared revision. A stale revision requires review.
Omit unchanged options to preserve inventory/role-default precedence.
The published API is in `scripts/docs/ADDON_API.md`; the complete new result contract
is in `scripts/docs/OPERATION_RESULTS.md`. Current support and qualification are in
`ADDON_SUPPORT.md` and `VALIDATION.md`. Never infer feature support from product version
alone. `capabilities.collection_baselines` declares collection floors that add-on installers/operators must satisfy; Core does not upgrade them. Unknown optional response fields may be ignored; unknown errors fail safely.
## Purposeful output: new in this candidate
Check `capabilities.operation_results`. Catalog `result` is null or a declaration with
schema, scope, required flag, limits and resolved `data_schema`. No new execute request
field is needed. Reports are available in summary and detail mode through the final
`RunResult.operation_result` and the existing final result event. There is no new raw
stdout event. Parse final JSON without scraping PLAY/TASK text or debug bodies.
A result contains `hosts[host]` with `status`, `schema`, `data`, `error`, plus optional
`global`. Render data only when status is `available`. Missing/withheld/invalid/null is
not an empty valid report. `complete` means all declared report slots arrived and passed
validation; it does not mean the operation succeeded. Partial native runs can contain
useful data from successful or failed targets. Preserve overall status/stage/exit,
`remote_work_may_have_started`, target summary and report availability as separate facts.
Never automatically retry a `result_validation` failure: remote changes may be complete.
Use a generic bounded JSON/table renderer for unknown schemas, with optional richer
views for known schemas. Escape text/HTML/Rich markup. Do not treat null versions as
zero, pending updates as installed, excluded services as failed starts, or redacted
configuration values as absent settings. Config reports expose sections but redact
recognized secret/command fields; restrict report visibility and retention accordingly.
## Safety and identity
Core checks active OS group membership and filesystem access. This is same-UID local
execution, not per-browser-user RBAC or an untrusted-tenant sandbox. The execution
account owns its 0600 keys and private staging. `service_user` is not a local UID switch.
`runtime.private_key_owner` affects new keys only. Do not relax permissions, export keys,
add generic sudo wrappers or run the web frontend as root to bypass the boundary.
Hardened service installers must provision both applicable staging locations (process
HOME and passwd home) and narrow ReadWritePaths exceptions. Retain ProtectHome/read-only,
ProtectSystem/strict and private tmp protections. Run `aimctl staging-check` inside the
actual sandbox. Core does not create accounts, modify units or restart add-on services.
Use OneRunCredentials or a private inherited FD, never passwords in request JSON, argv,
environment, URLs, logs, queue records or databases. Queue plans, not secrets. Fresh
credentials are required for each attempt. Password inputs are native defaults, not
Custom overrides or proof of password-only authentication. Cross-UID execution is not
provided by Core; independently managed executors must already have authorized access.
## Progress and final outcomes
`progress_mode: detail` is negotiated before preparation. Correlate play/task IDs,
handle withheld names and fixed error hints, and ignore anonymous progress to avoid
duplicate rows. Keep bounded queues and responsive sinks. Raw module/exception/debug
output remains unsupported. Purposeful data is a separate final schema, not an expansion
of host_result. Target outcomes come from Core's final native per-host stats, not task
counts. An interrupted/incomplete stream must not become partial success by inference.
## Ownership and handoff
Core and add-ons release independently. Core deployment preserves add-on code/state,
operator aim.yml, inventories, Vaults and keys; add-ons preserve Core in return.
Use ZIP + checksum + the deployer, not Git or patches. Coordinate active jobs before
source replacement. Only documented retired Core files are removed on upgrade.
The nine current report schemas are documented in OPERATION_RESULTS.md and advertised
by the catalog. This candidate has local regression coverage, not new managed-host
certification. Record independent acceptance under your actual identity/sandbox/runtime.
Request new schemas/capabilities from Core instead of adding private output workarounds.
## OS patch-wave presentation (3.3.0rc8)
Patch options remain catalog-driven. Interfaces should render the advertised reboot
message/delay and the Windows-only `os_patching_rescan_after_reboot` option. Its catalog
default is false; do not silently enable "fully patch" behavior in an add-on.
Windows submits the current selected categories as one native `win_updates` wave with
module-managed Windows Update sequencing and `reboot: false`. AIM evaluates reboot policy
after the wave returns; it does not expose or own a per-update scheduler. With post-reboot
continuation disabled, an AIM-performed reboot can still be a successful Core run while
`continuation_required: true` tells the operator another patch run is needed.
`remaining_updates_known: false` means the client must not invent or persist a next-wave
list. If a final read-only search was performed, `remaining_updates_known: true` makes
`pending` authoritative for that observation only.
Render `failed_updates[].reason/message/native_code_hex` instead of parsing fatal text or
Windows event logs. `install_not_allowed` deliberately does not mean "reboot required" by
itself. If Core reports `blocked_reason: preexisting_reboot_required`, that is the separate
preflight observation. Never auto-enable reboot, post-reboot continuation, or automatic job
replay.
+153
View File
@@ -0,0 +1,153 @@
# AIM public service API
**Core 3.3.0rc8; API/wire/event 1.0, stable additive 1.x; Ansible Core 2.19.11.**
Supported Python imports: `aim.services.v1` only. `aimctl` is a separate-process JSONL
interface in the AIM Python environment. No HTTP service, privilege broker or add-on
runtime is introduced. Core keeps its built-in terminal independently usable.
## Authorization and discovery
Protected calls check active primary/supplementary membership in configured
`required_group` plus normal local access. capabilities is unguarded non-sensitive
metadata. Caller-provided actor/user fields cannot authorize a request and are not
accepted. `--config` is a trusted operator path, never user input to a privileged wrapper.
| Operation | Python | Wire fields in addition to api_version/operation |
|---|---|---|
| capabilities | capabilities() | none |
| list_customers | list_customers() | none |
| list_hosts | list_hosts(customer) | customer |
| inventory_hierarchy | inventory_hierarchy(customer) | customer |
| list_playbooks | list_playbooks(customer) | customer |
| prepare | prepare(request) | request |
| readiness | readiness(request) | request |
| staging_check | staging_check() | none |
| execute | execute(request, expected_revision=..., credentials=..., event_sink=..., cancel=...) | request, expected_revision; secrets on separate FD |
Catalog metadata includes typed inputs and `result` (null or a resolved declaration
including data_schema). prepare returns normalized request, revision, credential
requirements/reasons, warnings, revision coverage and result_contract. Discovery and
prepare do not contact hosts or acquire secrets. readiness performs bounded local runtime
and staging checks; it does not prove remote authentication. staging_check is a bounded
local filesystem probe that may create missing private staging directories and deletes
its own test file, not operator contents. It invokes no Ansible or remote commands.
## RunRequest
```json
{"customer":"CUSTOMER","playbook":"debug_detect_host_roles","hosts":["HOST"],"overrides":{},"check":false,"key_mode":"none","become_password":false,"timeout_seconds":3600,"progress_mode":"summary"}
```
customer, playbook and hosts are required. Host count 1-1000; explicit unique inventory
names only, never arbitrary group/limit patterns. Core revalidates membership/platforms
and declared typed options. Omit options to inherit; 64 KiB maximum JSON overrides.
Secret-reference inputs are variable references, not passwords. Customer/host/catalog
code is trusted executable controller input, not an untrusted-tenant sandbox.
key_mode none preserves native inventory handling without importing the caller's agent.
customer loads only the canonical customer key into an owned agent: Linux target scope,
0600 key owned by execution UID, ssh-agent/ssh-add required. It never exports the key.
become_password true permits a default escalation password for a compatible catalog
scope; it does not independently turn on escalation. check uses native check mode, not
a guarantee of no effects for every plugin. timeout_seconds 10-86400 bounds the launched
playbook; preflight stages have their own deadlines. progress_mode summary/detail is
negotiated and part of the revision. Operation reports need no new request field.
## Lifecycle and revisions
Show prepared scope/options/key mode, result_contract and warnings before approval.
Check readiness in the real worker sandbox, obtain a slot, then fresh credentials.
execute re-prepares and checks revision before consuming secrets and before launch.
Known source/config/key bytes and metadata are hashed with limits, not a complete atomic
snapshot of dynamic includes, collections, external assets or arbitrary lookups. Stable
customer locks cover cooperating writers, not direct shell edits. Quiesce source updates.
`customer_vault_present` is conservative: a Vault can be needed by inventory credentials
even when catalog require_vault is false. Other reason codes include catalog_requires_vault,
catalog_requests_connection_password, request_requires_become_password and
encrypted_customer_key_selected. This is not exhaustive templated variable analysis.
## Private credentials
Accepted keys only: vault_password, connection_password, become_password,
ssh_key_passphrase. Literal single-line UTF-8 <=8192 bytes per value; no CR/LF/NUL;
nonempty Vault/connection/become passwords. No trimming, hashing or recursive templating.
Unencrypted keys need no passphrase. If selected encrypted key has no supplied phrase,
Core can read the literal vault_linux_ssh_key_passphrase after Vault unlock.
Python: use OneRunCredentials(values, ttl_seconds=120) and close in finally; TTL 1-300
seconds before consume. Custom providers implement bounded single-run consume(run_id).
Never prompt unattended or reuse a provider across jobs. Native executable password
sources use private same-UID sockets; helper files hold no secret material. No password
in argv/environment/request JSON/events. Native processes necessarily hold it in memory;
cleanup is not a memory-erasure guarantee. Supplied passwords are native defaults, not
forced Custom overrides or proof of password-only authentication.
Machine: one newline-terminated JSON object on stdin <=128 KiB within ten seconds.
Use `aimctl request --credentials-fd N` for execute, inheriting a private pipe/socket FD
>=3 with one EOF-terminated JSON object <=64 KiB. Close writer after sending. Ordinary
files/terminal/stdin/out/err FDs are rejected. Read deadline is five seconds when consumed.
Send request/secrets concurrently with event consumption; avoid filling a pipe before
starting the receiver. Pass FD using subprocess pass_fds, never shell text with passwords.
Cross-UID launcher transport is outside the Core-supported profile.
```bash
aimctl capabilities
printf '%s\n' '{"api_version":"1.0","operation":"list_customers"}' | aimctl request
printf '%s\n' '{"api_version":"1.0","operation":"prepare","request":{"customer":"CUSTOMER","playbook":"debug_detect_host_roles","hosts":["HOST"]}}' | aimctl request
```
## Events and final results
Each event has event_version, run_id, increasing sequence, UTC timestamp and kind.
Summary kinds stage/progress/stats/result stay. Detail adds negotiated play/task/host
metadata; see DETAILED_PROGRESS.md. Wire events are `{"type":"event","event":{...}}`.
Final response is `{"type":"response","api_version":"1.0","ok":true,"result":{...}}`
or ok:false with result/error. A final result event and final response are one job,
not two. Missing final response is unknown outcome, not success. Stdout is JSONL only.
RunResult fields: api_version, run_id, status (succeeded/failed/cancelled), stage,
nullable exit_code, remote_work_may_have_started, nullable error, aggregate task counts,
target_summary, targets, and operation_result (null when not declared). Target accounting
is in TARGET_OUTCOMES.md. Purposeful report availability, schemas, limits and errors are
in OPERATION_RESULTS.md. Native failures are preserved. A required report validation
failure after native exit 0 yields status failed/stage result_validation/exit_code 0;
native target stats remain unchanged. Never infer application data from ok/skipped counts.
Error fields: code, fixed message, stage, retryable, required_credentials. Handle unknown
codes safely, not by parsing message text. Pre-run ServiceError can occur without a
RunResult. No fake report or target success is synthesized for a rejected request.
| Error family | Action |
|---|---|
| access_denied / execution_disabled | Correct operator authorization/enablement, not bypass |
| invalid_request / invalid_target / invalid_options / unknown_playbook | Correct request, re-prepare |
| source_invalid / playbook_unavailable / source_symlink_unsupported | Repair trusted source |
| resource_busy / review_stale | Wait or re-review; retain no stale secrets |
| runtime_missing / runtime_version_unsupported / collections_missing | Repair approved native runtime |
| controller_staging_* | Repair scoped worker staging inside sandbox; no password loop |
| runtime_credential_defaults_unsupported | Resolve unsupported global Vault sources explicitly |
| credential_required / credentials_expired / invalid_credential_channel | Fresh bounded provider/channel |
| vault_unlock_failed / key_load_failed | Correct secret/access/format before launch |
| syntax_check_failed | Authorized native diagnostics, no raw fallback |
| host_unreachable / playbook_failed | Inspect outcomes; no automatic replay |
| operation_result_* | Inspect report availability; remote work may be finished; never replay automatically |
| invalid_event_stream / event_bridge_* / event_limit | Incomplete stream is not success |
| timeout / cancelled / event_sink_failed / internal_error | Explicit recovery decision, preserve remote-work flag |
Cancel with threading.Event in Python or SIGINT/SIGTERM to aimctl. Core stops owned
process groups/agents, not completed remote changes or asynchronous external work.
Use a bounded nonblocking event sink; slow browsers must not stall execution.
## Runtime and detailed contracts
The configured ansible-playbook plus sibling Vault/Galaxy must report canonical 2.19.11
and consistent native Python/module identity. Collections, roles, local staging, remote
Python/PowerShell, HOME/known-hosts and service restrictions need actual acceptance.
The service selects root_dir/ansible.cfg or its fallback; it does not inherit arbitrary
shell environment/credential/debug settings. Supported collection search path is explicit.
No auto-install, recursive ownership repair, remote_tmp override or service-unit edit.
See EXECUTOR_STAGING.md (both applicable homes), INVENTORY_HIERARCHY.md, TARGET_OUTCOMES.md,
DETAILED_PROGRESS.md and OPERATION_RESULTS.md. ADDON_SUPPORT.md lists deliberate exclusions.
VALIDATION.md records present evidence; SANITY.md is the only current acceptance checklist.
+49
View File
@@ -0,0 +1,49 @@
# AIM 3.3.0rc8 support matrix
Service/wire/event 1.0 is stable and additive. Implementation status and controller
qualification are different. Use live capabilities; the JSON file alongside this guide
is only a default capability snapshot. [VALIDATION.md](VALIDATION.md) is the sole current
evidence record.
| Capability | Implementation | Activation/qualification |
|---|---|---|
| Discovery/catalog/explicit prepare | Supported | Required OS group/filesystem access |
| Inventory hierarchy v1 | Supported, read-only | No independent inventory semantics in add-ons |
| Staging preflight v2 | Supported | Test both applicable homes inside the real sandbox |
| Native execution | Supported same-UID profile | Explicit config opt-in; ansible-core 2.19.11 and ansible.windows >=3.8.0,<4.0.0 |
| Private credentials | Supported one-run provider/FD | Not generic request JSON; no native prompt automation |
| Detailed progress v1 | Supported opt-in | Canonical callback/target acceptance required |
| Per-target outcomes v1 | Supported summary/detail | Native final stats, not application-completion inference |
| Purposeful operation reports v1 | Implemented candidate | Nine per-host schemas; native acceptance pending |
| Global operation report transport | Implemented candidate | No shipped global publisher; separate native qualification |
| Parsed Checkmk user config | Implemented candidate | Restricted basename/size; recognized secrets and commands redacted |
| Custom credentials / password-only verification | Unsupported | Supplied passwords remain native defaults |
| Cross-UID broker, key export, automatic owner migration | Unsupported | Independent execution identity must already be authorized |
| API inventory/Vault mutation or bootstrap | Unsupported | Existing terminal workflows remain available |
| Arbitrary debug/stdout/stderr/results | Unsupported | Only explicitly declared validated report data |
| Automatic replay after launch | Unsupported | Explicit new operator decision required |
A structured result is operational data and can reveal host inventory/configuration
choices even after credential redaction. Add-ons must restrict access, safely render
text and bound retention. Core does not implement their browser RBAC/CSRF/session policy.
## OS patch presentation in 3.3.0rc8
Add-ons obtain reboot delay/message and Windows post-reboot continuation from normal catalog
metadata; no private UI contract is required. The `os_patching_rescan_after_reboot` default
is false. A normal Windows patch run submits one selected update wave to native
`ansible.windows.win_updates` with `reboot: false`, then applies AIM reboot policy after the
wave returns. Show `continuation_required: true` as a request for
a new operator-approved run, not as a Core failure and not as permission to submit another
job automatically.
If `remaining_updates_known` is false, do not infer the post-reboot patch state from the
pre-reboot queue. If true, `pending` is a read-only final discovery. For Windows Update
failures, render the bounded `failed_updates` fields; do not parse raw task/event text.
`install_not_allowed` can mean an active installer or mandatory reboot, while
`preexisting_reboot_required` is a distinct AIM preflight result.
## Veeam O365 plugin relocation
The Checkmk Veeam O365 script destination changed in Core 3.3.0rc8, but this is not an add-on API change. Consumers continue to use Core catalog/execution/results and must not reconstruct Checkmk filesystem paths independently.
+99
View File
@@ -0,0 +1,99 @@
# AIM authoritative development guidelines
**Baseline: 3.3.0rc8.** This is the current Core guide. Operator instructions take
precedence over earlier project documentation. Keep the built-in terminal in AIM.
## Product and release ownership
Core owns domain/execution truth, storage, reusable services and both first-party
interfaces. Dependencies point interface -> services/domain/runtime; services must
not import `aim.ui` or any add-on. Add-ons consume only `aim.services.v1` or `aimctl`.
No Core change may require inspecting or editing an independent add-on's codebase.
Release complete source ZIPs with SHA-256 sidecars and the operational deployer. No
Git/patch dependency, `.patch` file, release manifest, bundled test/validator machinery,
cache, inventory, Vault, private key, environment, downloaded package or collection.
Do not mutate a published candidate. Every source change updates the product version
and CHANGELOG. Service/wire/event 1.0 evolves additively; breaking changes require a
separately versioned contract. Product and API versions are not interchangeable.
## Layout and preservation
`/etc/ansible/scripts` contains source, pyproject, global `aim.yml`, docs and the WinRM
helper. `playbooks` contains the catalog, schemas and report filter plugins. `roles`
contains execution roles. Customer `.aim.yml` remains in its customer directory.
Deployment preserves operator config, inventory/Vault/keys, add-ons, environments and
customer assets. Preview first; quiesce writers/runners; maintain source/launcher
recovery. Retire only explicitly named old Core documents, never arbitrary Markdown.
Do not reset LDAP groups, permissions or ownership recursively. Stable lock inodes and
metadata-preserving atomic replacement apply. Source rollback cannot undo remote work.
## Execution and credentials
Canonical native runtime: ansible-core 2.19.11. AIM and add-on Python environments may
be separate; do not install/upgrade Ansible during discovery or deployment. New service
profiles stay opt-in. Require the actual execution UID/groups, owner-only private keys,
usable collections and sandbox staging. Check both process/config HOME and passwd/NSS
home paths where local delegation expands `~account`. Never override remote_tmp globally,
switch UIDs or edit service units to hide a staging/access failure.
Credentials use bounded one-run providers/private descriptors/sockets, never request
JSON, argv, environment, output, examples or persisted plans. Supplied passwords remain
native defaults, not forced inventory overrides. Existing Vault unlock retries are
pre-launch only. Preserve selection on typos and Ctrl+C cancellation; never replay a
launched editor/playbook automatically. Source and prepared revisions are change
detectors, not authorization tokens or isolation from uncooperative shell edits.
## Structured purposeful results
Use `ansible.builtin.set_stats` with `data.aim_output`, `protocol: aim_output_v1`, a static
schema ID, `per_host` matching the declaration and `aggregate: false`. Exactly one
publication per host (or one global publication) per execution. Result normalization
belongs in runbook roles/filter plugins, not hard-coded playbook branches in Core.
Declare result shape/scope/required/sensitivity in the catalog. Define bounded types in
`playbooks/schemas/`. The generic validator owns type/field/size/lifecycle enforcement;
new schemas require no Core dispatch additions. Revalidate at the callback and public
boundary. Preserve no_log provenance: final custom stats alone are not authorization
to expose data. Unknown custom stats/debug/module results remain private.
Operation reports, target outcomes and progress are separate. A required missing or
invalid report can fail result validation after native exit 0; preserve exit_code and
native target stats, and never replay. Interrupted output is indeterminate. On native
failure retain available final reports without replacing the original failure.
Declare operational field meanings precisely: bytes, observed versions, nullable
unknowns, bounded fixed error codes. Never guess an installed version from a staged
filename, imply a start return means a service stayed running, or call check-mode
predictions completed changes. Patch reports must not silently truncate update lists.
Checkmk config is parsed and redacted before publication; the filename alone does
not make passphrases/MRPE command credentials safe. Raw configuration diffs stay private.
Windows patching is operator-wave controlled: let the supported `win_updates` module process the currently approved update wave with module reboot disabled, then treat any resulting reboot as the boundary. Post-reboot rediscovery/continuation must be an explicit catalog option defaulting false; never turn a normal patch run into an implicit "keep patching until current" workflow. Preserve bounded per-update HRESULT reporting and never equate `0x80240016` alone with a proven pending reboot.
## Existing compatibility and Checkmk boundaries
Keep HTTPS WinRM 5986 and guarded legacy certificate fallback. Preserve standalone
local-account token handling. Do not add legacy-only external bootstrap dependencies.
Windows bootstrap compatibility is separate from collection/module requirements;
never claim every new playbook was tested on Server 2012 R2/PowerShell 4.
AIM owns catalogued check filenames, not whole directories. Unknown files stay.
Windows configuration management owns only `plugins:`; preserve `local`, `mrpe`,
`global`, unknown sections and comments. Keep the first-line notice plus section
ownership marker. Unmanaged sections are not empty generated mappings. Avoid unmatched
quote characters even in inline free-form shell comments. No unrelated firewall policy
changes or cleanup expansions in a reporting release.
AIM-managed persistent Windows Checkmk files use the shared `checkmk_windows_acl` role: enable parent inheritance per file and guarantee well-known SID access for SYSTEM/local Administrators (FullControl) and both application-package principals (ReadAndExecute). Do not add customer-specific administrator/user ACEs, recurse over whole Checkmk directories, or rewrite unknown/operator-file ACLs.
## Verification and documentation
Tests live outside the shipped release. Parse/compile sources, validate YAML/Jinja,
exercise credential/process/output boundaries and installer rollback. Use native
Ansible/collections when available and label simulations distinctly. Record failures,
missing dependencies and platform acceptance honestly in `scripts/docs/VALIDATION.md`.
Keep exactly one current API, release notes, handoff, sanity checklist and validation
record; the documentation index names the authority for each topic. Historical
CHANGELOG entries are retained, not represented as present-day capability guarantees.
+268
View File
@@ -0,0 +1,268 @@
# Changelog
## 3.3.0rc8
- Normalize ACLs on persistent AIM-managed Windows Checkmk scripts and `check_mk.user.yml` after creation/update.
- Re-enable parent ACL inheritance per managed file and guarantee locale-independent well-known principals by SID: SYSTEM and local Administrators with FullControl; ALL APPLICATION PACKAGES and ALL RESTRICTED APPLICATION PACKAGES with ReadAndExecute.
- Do not add customer-specific administrator/user ACEs and do not recursively rewrite Checkmk directories or unknown/operator files. Existing intentional parent/explicit ACEs are not blindly purged.
- Add the reusable `checkmk_windows_acl` role and apply it only to the exact AIM-managed persistent file paths.
- Keep Checkmk script placement, structured-result contracts, service/wire/event API 1.0, and Ansible Core 2.19.11 unchanged.
## 3.3.0rc7
- Move `veeam_o365_status.ps1` from the Windows Checkmk local-check directory to `C:\ProgramData\checkmk\agent\plugins`.
- If an earlier rc7 build placed that file under the Checkmk built-in plugin directory, selected deployment removes that known stale copy after the custom plugin is staged.
- Add the exact `$CUSTOM_PLUGINS_PATH$\veeam_o365_status.ps1` execution rule when Veeam VBO is detected, ahead of the built-in plugin deny catch-all.
- Remove the legacy AIM-managed local copy after selected deployment and teach explicit cleanup to remove the current custom-plugin path plus both known legacy locations without touching unknown files.
- Keep Checkmk public structured-result schemas and service/wire/event API 1.0 unchanged.
- Replaced the Windows disk-usage PowerShell/Get-PSDrive collector with `community.windows.win_disk_facts`; Windows capacity reports now describe attached local volumes and intentionally exclude mapped/network drives.
- Kept the public `filesystem_usage_v1` result schema stable while normalizing native disk/partition/volume facts into the existing fields.
- Audited the complete playbook/role tree for native Ansible module coverage.
- Replaced custom Windows pending-reboot registry probing with `ansible.windows.win_reboot_info` and added bounded reboot-source reporting.
- Raised the supported ansible.windows baseline to >=3.8.0,<4.0.0 and added explicit readiness/terminal version checks.
- Replaced Checkmk Windows config file shell reads with `win_stat` + `slurp`.
- Replaced Linux package snapshot/version commands with `package_facts` and Linux Checkmk final service-state command reporting with `service_facts`.
- Documented reviewed custom command/PowerShell/raw call sites that remain because native modules do not preserve the required behavior.
## 3.3.0rc3
- Change Windows OS patching from one broad `win_updates` install request to an explicit discovery queue with one update installed per invocation. The queue is deterministic and places drivers last.
- Add Windows-only catalog option `os_patching_rescan_after_reboot` (`false` by default). A patch-triggered reboot ends the run by default; post-reboot discovery/install continuation requires explicit operator opt-in.
- Preserve one operator-approved patch wave per run by default. A completed reboot boundary reports `continuation_required: true` and does not perform a post-reboot search unless continuation was enabled.
- After a wave completes without reboot, perform one read-only final discovery for reporting only; do not append newly applicable updates to the already-approved install queue.
- Expand `patch_summary_v1` with Windows patch-cycle/continuation facts and bounded per-update HRESULT classification. `0x80240016` is reported as `install_not_allowed`, not treated as proof of a reboot.
- Stop the Windows wave on an individual update failure, publish already completed updates plus the bounded failure, and never replay the update/job automatically.
- Keep reboot notification/delay, Linux patching, service/wire/event API 1.0, Ansible Core 2.19.11 and the generic operation-result transport unchanged.
- Native Windows controller acceptance of the sequential queue and post-reboot continuation remains required before stable promotion.
## 3.3.0rc2
- Improve OS patch reboot handling without changing Core API/wire/event 1.0: expose a cross-platform reboot delay (minutes) and reboot notification message through catalog metadata.
- Windows patching now disables win_updates implicit reboot and uses an explicit win_reboot so the configured message/delay is honored; Linux uses the corresponding reboot module message/delay semantics.
- Detect a pre-existing pending reboot before starting new patch work where the platform provides a supported signal. If automatic reboot is disabled, publish a structured blocked patch result and fail with an explicit reboot-required message instead of a generic later fatal.
- When a patch run newly requires reboot and automatic reboot is disabled, keep the successful update run successful but report reboot_required/reboot_deferred in patch_summary_v1. A later run remains blocked until the host is rebooted.
- Extend patch_summary_v1 with pre/post/deferred reboot facts, configured delay and a typed preexisting_reboot_required block reason; package/update reporting remains unchanged.
- Add the current fresh-install INSTALLATION.md to the consolidated documentation set.
- Native Windows/Linux controller acceptance for these rc2 patching changes remains pending; see VALIDATION.md and SANITY.md.
## 3.2.1rc2 - inventory hierarchy, split-home staging and per-target outcomes
- Adds read-only `inventory_hierarchy` discovery with nested parent/subgroup relationships and direct host membership. The public tree exposes group names/paths and host names only; variables remain private and execution still requires explicit reviewed host names.
- Corrects controller staging preflight for hardened executors whose process `$HOME` differs from the execution account passwd home. Core now probes controller `local_tmp` using the process/config home and the default delegated `connection: local` POSIX temp path using the passwd/NSS account home. No home, permission or service-unit mutation is performed.
- Adds additive `target_outcome_summary_v1` data to every execution `RunResult`. Overall Core status/exit semantics remain unchanged; consumers receive requested-target accounting plus per-target `successful`, `failed`, `unreachable`, `not_started` or `indeterminate` outcomes derived from native final host stats.
- Target summaries are available in both summary and detail progress modes and do not require consumers to reconstruct task events. Incomplete/cancelled execution without final stats never fabricates target success.
- Keeps service/wire/event API 1.0, canonical Ansible Core 2.19.11, credentials, playbooks, roles, WinRM and remote execution semantics unchanged. This is an additive discovery/readiness/result candidate.
## 3.2.1rc1 - controller staging preflight and hardened-executor default
- Makes a writable, private controller staging directory an always-checked service requirement. Readiness and execution validate staging before credentials/native runtime inspection, with a second check before playbook launch. Failures return fixed readiness errors rather than late delegated-task unreachable results.
- Adds protected `staging_check` / `AimService.staging_check()` and `aimctl staging-check`, plus additive capability/readiness fields. No credentials, customer inventory or native Ansible command is needed for the local probe.
- Performs real create/write/flush/read/remove operations under the current UID and sandbox, bounded to ten seconds. Missing directories are privately created; existing permissions, owners and contents are never repaired or purged.
- Preserves native controller temporary configuration and remote-host settings. Checks the default POSIX local-connection staging path; explicitly warns when global remote_tmp overrides or inventory-specific settings are outside probe coverage.
- Establishes the generic add-on deployment default: provision private staging, append a directory-specific ReadWritePaths exception and an in-sandbox startup check while retaining ProtectHome/ProtectSystem/PrivateTmp/NoNewPrivileges. Core never edits independent service units or add-on code.
- Records operator confirmation that the narrow staging exception allowed the Checkmk service job to install successfully on a Windows 11 client. New candidate acceptance and detailed-progress/other-target qualification remain separate.
- Keeps API/wire/event 1.0, Ansible Core 2.19.11, existing roles/playbooks, WinRM, credentials, terminal behavior and deployment logic unchanged. Full replacement ZIP plus checksum; no patches or bundled validators.
## 3.2.0 - safe detailed execution progress
- Adds optional `progress_mode: detail` to the stable service/wire/event 1.0 contract. The default `summary` mode retains anonymous progress and existing consumers. Capabilities advertise `play_task_host_v1`; clients opt in before requesting it.
- Detailed events expose plays, no-host skips, tasks/handlers, per-host outcomes, retry/poll notices and recaps, with run/play/task correlation. Native first-party terminal output remains unchanged.
- Uses bounded unexpanded source labels, explicit reviewed host names and fixed diagnostic hints. Suppresses unsafe/template/no_log labels and protected result details; never forwards raw stdout/stderr, module arguments, variable/config dumps, exception fields or loop-item payloads.
- Revalidates native frames, sequences and recap totals before public emission. Missing, incomplete, malformed or over-limit detailed progress cannot silently become a successful result. No automatic replay or rollback of remote work is introduced.
- Keeps Ansible Core 2.19.11 canonical, external execution disabled by default, and existing credential/ownership/deployment policies. No playbook, role, WinRM helper, inventory, key, add-on or deployment behavior is changed.
- Updates release notes, additive API contract, support metadata, renderer guidance and the controller acceptance checklist. Local tests use simulated Ansible callbacks/CLIs; new detailed execution still requires native controller/target acceptance.
## 3.1.0 - accepted service-v1 baseline
- Promotes the controller-accepted 3.0.0rc21 candidate to the stable 3.1.0 minor release without changing playbook, role, credential, WinRM, Checkmk, deployment, or service behavior.
- Freezes AIM service/wire/event API **1.0** for additive 1.x evolution. AIM continues to ship the built-in terminal plus `aimctl`; add-ons consume only the documented service boundary.
- Records real-controller acceptance of Windows 11 terminal management, Vault retry behavior, group/subgroup select-all, launcher deployment, `aimctl` discovery/preparation, and native Ansible Core 2.19.11 readiness.
- External non-interactive execution remains disabled by default and requires separate acceptance by the chosen add-on/runtime identity. Unsupported boundaries remain unchanged: cross-UID execution, forced Custom password override, private-key export, API inventory/Vault mutation, raw sensitive task output, and automatic replay after remote work may have started.
- Release delivery remains a complete replacement ZIP plus SHA-256 sidecar; no Git, `.patch`, release manifest, or bundled development validator is used.
## 3.0.0rc21 - launcher symlink deployment hotfix
- Fixes the rc20 ZIP deployer rejecting an existing `/usr/local/bin/aim` symlink before it could refresh the installed `aim` and `aimctl` launchers.
- Allows only recognized existing AIM launcher symlinks at the final launcher path; symlinked parent directories, broken links, non-regular targets, and unrecognized launcher contents remain rejected.
- Deployment now backs up the symlink itself, atomically replaces it with the managed launcher, verifies `aim` and `aimctl`, and restores the original symlink target during rollback.
- No service API, terminal workflow, playbook, role, WinRM, Checkmk, credential, or remote-execution behavior changes from rc20. Service API / wire / event remain 1.0 and Ansible Core 2.19.11 remains canonical.
- Records the operator-observed rc20 deployment failure as the regression target. Local disposable testing passed dry-run, symlink replacement, installed command verification, and rollback restoration.
## 3.0.0rc20 - terminal stabilization and stable services v1 handoff
- Adds interface-supplied, operation-scoped Vault unlock/retry before terminal Vault edits, reads and normal catalog playbook launches. A rejected password keeps the selected workflow. Empty/invalid input re-prompts; Ctrl+C cancels. Only the native Vault decryption rejection from a bounded local `ansible-vault view` preflight is retried, never an editor or launched playbook.
- Reuses the private same-UID password-client channel for the validated operation: no password in argv/environment/helper contents and no second Vault prompt after successful prevalidation. Terminal Vault-password input preserves literal whitespace. Legacy internal manager callers without terminal interaction retain native prompts. Additional/inline Vault files are not exhaustively prevalidated; late failures are never automatically replayed. Vault create/new-encryption prompts remain native.
- Adds "Select all hosts in this group/subgroup" versus individual selection after group selection. Parent scopes include descendant subgroups; overlapping memberships are deduplicated. Execution still uses explicit host names and the existing final run review/confirmation.
- ZIP deployment now installs/refreshes and smoke-tests both `aim` and `aimctl` launchers with the existing AIM Python, without pip, dependency installation or changes to the Ansible environment. Adds `--aim-python` and `--bin-dir`, conservative interpreter discovery, previewed launcher changes and launcher recovery alongside source recovery. Unknown commands/symlinks are refused. A separate target requires its own explicit command directory.
- Returns safe `invalid_target` and `invalid_options` service errors for known validation failures; actual inaccessible/invalid source remains `source_invalid`. Error-code fallback remains required for v1 clients.
- Adds optional `credential_requirement_reasons` to prepared results and advertises `credential_requirements_policy`. A present conventional customer Vault remains conservatively required even with `require_vault: false`; this is not a claim of full effective-variable credential analysis.
- Freezes the documented service/wire/event 1.0 contract as the supported additive 1.x boundary. AIM keeps its built-in terminal and machine client; internal managers remain private. No add-on source or interface framework is required.
- Records operator-reported rc19 Windows 11 terminal onboarding/service-account authentication, targeting, interruption, metadata/prepare and native 2.19.11 readiness results separately from rc20 local regression evidence and pending controller re-tests. External execution remains opt-in/disabled by default; neither readiness nor CLI success proves the non-interactive execution path.
- Updates add-on guidelines, API/support contract, deployment instructions, `RC20_SANITY_TESTS.md` and `RC20_HANDOFF.md`. No WinRM/Checkmk/other remote playbook or role behavior is changed. No Git/patches/release manifest/development validators are shipped.
## 3.0.0rc19 - independent add-on service foundation and ZIP deployment
- Adds `aim.services.v1`, `aimctl` and a source-tree machine wrapper. Generic API/wire/event version 1.0; no add-on/web-framework imports or need to inspect an add-on codebase.
- Adds capabilities, authorized metadata reads, typed explicit-host/catalog preparation, bounded known-source revisions, local readiness and opt-in same-UID native Ansible execution. Shared target validation now protects direct CLI manager calls as well as the UI.
- Establishes **ansible-core 2.19.11** as the canonical supported runtime, separately from the AIM/add-on Python environments. Native sibling CLI identity and required collections are checked without installation side effects.
- Adds private one-run credential providers, inherited credential-FD support, native executable password sources and safe progress/results. Keeps native inventory/Vault precedence: supplied connection/become passwords are defaults, not a forced Custom override. Existing terminal `@prompt` behavior is retained; `--ask-vault-pass` is also interactive and does not solve unattended input.
- Adds bounded Vault preflight, owned per-run SSH agents/process groups, cancellation and conservative remote-work reporting. Unencrypted keys no longer require a Vault passphrase. CLI agent probing rejects stale sockets and no longer transports key passphrases through environment values.
- Preserves owner/group/mode/attributes before atomic replacements, including inventory/configuration/Vault publication and restore paths. Existing non-root shared-group writers retain their legacy owner-transfer exception with a warning; full managed single-writer ownership is not claimed.
- Makes cooperative customer locks persistent and reentrant instead of unlinking their inode. Extends locking across key/Vault/config mutations and playbook execution. Direct shell edits and arbitrary external writers remain outside advisory-lock guarantees.
- Adds optional `runtime.private_key_owner` for new keys only, with exclusive staged keypair publication. Does not rename keys, change remote `service_user`, re-own existing keys, switch UIDs or migrate permissions automatically.
- Adds the operator-approved standard-library `deploy/deploy.py`, dry-run-first source replacement, explicit quiescence and protected source recovery. Delivers a ZIP plus SHA-256 sidecar without Git, patches, release manifests or bundled development validators. Existing aim.yml, add-ons/runtime state, inventories/Vaults/keys, environments and unknown customer files are preserved.
- Publishes `ADDON_AGENTS.md`, `ADDON_API.md`, `ADDON_SUPPORT.md`, `addon-support-v1.json`, and `RC19_HANDOFF.md`. Exposes unsupported capabilities explicitly: Custom overrides, cross-UID launchers, key export, API mutations and raw sensitive task results.
- External execution defaults to disabled. Development validation includes real Python/Unix IPC/filesystem/process tests and simulated native CLI contract fixtures, **not** live Ansible 2.19.11, SSH or WinRM acceptance. See the handoff for evidence and controller gates.
- No Checkmk task/template, WinRM bootstrap or remote configuration behavior was changed. The earlier question about a remote Checkmk configuration backup remains separate and unimplemented.
## 3.0.0rc18 - Checkmk Windows config parser hotfix
- Fixes the `checkmk_configure_agent` Windows plugins-section update task failing during Ansible argument parsing before reaching the managed host.
- Removes an unmatched apostrophe from an embedded `win_shell` PowerShell comment; Ansible free-form shell argument parsing still scans quote characters inside the script text, including comments.
- Does not change Checkmk configuration behavior: AIM still replaces only the marked top-level `plugins:` section and preserves all other `check_mk.user.yml` content.
## 3.0.0rc17 - read-only Windows Checkmk configuration inspection
- Adds `checkmk_read_windows_config.yml`, a read-only Windows playbook that displays the current `check_mk.user.yml` without changing the host.
- Reports the resolved configuration path, file size, UTC last-write timestamp, and complete current file contents; a missing file is reported without failing or creating it.
- Adds an optional `checkmk_windows_user_cfg` path override for nonstandard agent layouts while retaining `C:\ProgramData\checkmk\agent\check_mk.user.yml` as the default.
- Exposes the operation in the AIM Checkmk catalog and documents it as a safe pre/post-rollout inspection tool.
- Uses the standard Windows shell path rather than `win_powershell`, keeping this inspection operation usable on legacy PowerShell 4 hosts such as Windows Server 2012 R2.
## 3.0.0rc16 - section-scoped Checkmk user configuration
- Changes Windows `check_mk.user.yml` handling from full-file rendering to section-scoped editing. AIM now replaces or appends only the top-level `plugins:` section.
- Preserves existing `global`, `winperf`, `fileinfo`, `logwatch`, `local`, `mrpe`, unknown top-level sections, and unrelated comments instead of resetting them to AIM defaults.
- Keeps an AIM ownership notice on the first line and adds a dedicated ownership comment immediately before the AIM-managed `plugins:` section so the management boundary is explicit.
- Removes the rc15 `checkmk_manage_local_execution` / `checkmk_extra_local_patterns` rollout controls: local-check execution is no longer managed by this role at all. Existing local execution policy is preserved.
- Preserves existing MRPE configuration instead of writing `mrpe.config: []`.
- Retains the existing AIM plugin execution ordering and role-based plugin rules; unknown local/plugin files remain untouched.
- Existing hosts already overwritten by an earlier full-file rollout cannot have lost settings reconstructed automatically; restore those settings from the host's previous configuration/backup if needed, then rc16 will preserve them on subsequent runs.
## 3.0.0rc15 - Checkmk local execution inheritance hotfix
- Fixes Windows Checkmk user configuration generation so `checkmk_manage_local_execution: false` omits the `local` section instead of writing `local: {}`.
- This preserves the effective local-check execution policy inherited from Checkmk default/bakery configuration and prevents custom files in `C:\ProgramData\checkmk\agent\local` from becoming non-executable merely because AIM rendered the user configuration.
- When `checkmk_manage_local_execution: true`, AIM now explicitly writes `local.enabled: true` together with the managed `local.execution` rules.
- Does not change Checkmk script-file ownership: unknown local/plugin files remain untouched, while known AIM-managed filenames may still be replaced.
- The Checkmk user configuration file remains fully rendered by AIM in this hotfix; merge-preserving ownership of arbitrary pre-existing user-config keys is a separate design change.
## 3.0.0rc14 - clean replacement bundle layout
- Changed release packaging to a complete replacement/fresh-install model instead of shipping patch/merge deployment mechanics.
- Removed `release-manifest.json`, release updater/deployer tooling, bundled release-validation tooling, historical migration guides, `.patch` artifacts, and generated Python caches from the distributable archive.
- Moved the controller-wide AIM configuration from `/etc/ansible/aim.yml` to `/etc/ansible/scripts/aim.yml` and included a clean default `scripts/aim.yml` in the bundle.
- Moved the maintained one-time WinRM bootstrap helper to `scripts/AIM-WinRM-OneTime.ps1`.
- Retained operational documentation, changelog, playbooks, roles, collection requirements, and the complete AIM application source required for a new installation.
## 3.0.0rc13 - forgiving WinRM password prompts
- Temporary/bootstrap WinRM credentials are validated before the requested Windows access operation begins.
- A failed WinRM credential probe now re-prompts for the password instead of consuming/skipping the host operation. Ctrl+C still cancels normally.
- Empty temporary passwords are re-prompted instead of failing the workflow.
- Service-account password entry now loops on empty values or confirmation mismatches instead of aborting back to the menu.
- Introduces a dedicated `WinRMConnectionFailed` error so only an actual WinRM probe failure triggers credential re-entry; missing commands and unrelated controller errors still fail normally.
## 3.0.0rc12 - Windows Server 2012 R2 WinRM certificate hotfix
- Fixes `New-SelfSignedCertificate` on Windows Server 2012 R2 / PowerShell 4 systems that expose `-CertStoreLocation` but reject `Cert:\LocalMachine\My` with `InvalidStorePathException`.
- Certificate creation now retries only that specific failure from inside `Cert:\LocalMachine\My`, preserving normal behavior and error reporting on newer Windows versions.
- Keeps capability detection for optional `-FriendlyName` and `-TextExtension` parameters and applies the friendly name after creation when required.
- Adds `scripts/tools/AIM-WinRM-OneTime.ps1` as the maintained standalone one-time WinRM HTTPS bootstrap script using the same compatibility logic.
- Adds top-level `AGENTS.md` as the authoritative development/release/compatibility guideline for continued AIM development.
## 3.0.0rc11 - Windows local host_vars credential-model repair hotfix
- `Prepare local account` now inspects every selected host's `host_vars/<fqdn>/main.yml` before prompting for credentials.
- Empty legacy host-vars files and partial overrides are explicitly identified as incomplete when `ansible_user` or `ansible_password` is missing/blank.
- The operator chooses the desired local credential model: shared-local, host-specific local, or retain complete existing overrides.
- `Retain existing` is rejected while any selected host has incomplete credential overrides, preventing a successful account bootstrap from leaving an unusable empty/partial host-vars file behind.
- After the local account bootstrap and independent WinRM verification succeed, AIM reapplies the selected shared-local or host-specific model to `host_vars`, preserving unrelated custom variables/comments.
## 3.0.0rc10 - legacy inventory consolidation/template comments hotfix
- Template consolidation now creates missing platform `group_vars/<platform>/` directories and `main.yml` files, matching the existing preview message instead of silently skipping absent parent directories.
- Existing group-vars values/custom variables remain non-destructively preserved; malformed YAML continues to be skipped rather than repaired implicitly.
- Windows shared-local and host-specific credential overrides now write a clear AIM ownership/context comment as the first line of `host_vars/<host>/main.yml`.
- Template validation/consolidation also detects older AIM local-credential `host_vars/<host>/main.yml` files that lack that header and adds it without changing existing credential values or custom variables.
- Switching a host back to the domain credential model removes only AIM's known local-credential header while preserving unrelated custom comments/keys.
## 3.0.0rc9 - Standalone Windows WinRM local-admin hotfix
- `Prepare local account` now idempotently sets `LocalAccountTokenFilterPolicy=1` before the service-account WinRM verification.
- Keeps the existing bare `svc_bf-ansible` username behavior because that is known to work once remote UAC token filtering is disabled.
- Domain-account and domain-GPO behavior are unchanged.
- The GPO WinRM payload retains direct `Get-NetIPAddress` discovery and optional additional SANs for Hyper-V/complex networking cases.
## 3.0.0rc8 - Consolidation permission handling hotfix
- Treat existing shared inventory mode `0660` as already correct and avoid redundant chmod attempts.
- Template consolidation uses best-effort permission normalization after a successful content update.
- `EPERM`/`EACCES` during consolidation permission normalization is reported as a warning instead of failing the operation.
- Security-sensitive permission operations outside template consolidation remain strict.
## 3.0.0rc7 - WinRM certificate compatibility hotfix
- Makes the generated `AIM-WinRM-Setup.ps1` compatible with PKI cmdlet implementations that do not expose `New-SelfSignedCertificate -TextExtension` and/or `-FriendlyName`.
- Builds the certificate call from the parameters actually supported by the target host. The base `-DnsName` + `CertStoreLocation` path remains an SSL server certificate; the explicit Server Authentication EKU extension is added when `TextExtension` is available.
- Applies the `WinRM` friendly name after creation when it could not be supplied as a cmdlet parameter, preserving AIM's existing certificate-reuse behavior.
- Does not change listener, firewall, GPO link, scheduled-task timing, service-account or access-policy behavior.
## 3.0.0rc6 - coordinated playbook/role migration candidate
- Adds top-level `Multi-customer operations` next to `Open customer`; it is never launched from inside a customer context.
- Adds multi-customer catalog playbook execution with per-customer `hosts.yml` target selection, shared typed run options and one isolated Ansible invocation per customer.
- Runs selected customers sequentially, keeps native Ansible/Vault/password prompts live, continues with the next customer after an ordinary playbook failure, and stops the remaining batch when the operator interrupts the current command.
- Adds clear per-customer live-output separators and a final summary with target counts, result and duration, plus failure/skip details.
- Keeps customer-specific inventories, Vault IDs, SSH preparation and target limits isolated; inventories are never merged.
- Changes multi-select `a` into a filter-aware toggle: when all currently matching items are selected it clears them; otherwise it selects all currently matching items. Selections outside the active filter are unchanged.
## 3.0.0rc5 - coordinated playbook/role migration candidate
- Stores the controller-wide Checkmk agent source as structured settings in `/etc/ansible/aim.yml`: protocol, base URL/host, site name, version, patchlevel and revision. AIM always appends `/check_mk/agents` and builds the package version as `<version>p<patchlevel>-<revision>`. Legacy complete URL/version settings are migrated in memory and normalized on the next save.
- Moves `Save settings` to the bottom of the Settings menu.
- Adds `a` to every multi-select selector to select all items matching the current filter (across all filtered pages).
## 3.0.0rc4 - coordinated playbook/role migration candidate
- Expanded `debug_show_disk_usage` to Windows and Linux. Windows reports all filesystem drives visible to the WinRM session; Linux reports common operational mounts and network/storage filesystems while excluding pseudo/system mounts.
Breaking source layout / entrypoint changes from AIM 2.2.2. AIM and its catalog,
playbooks and roles must be installed together. No legacy wrappers are shipped.
- Introduces `playbooks/aim_catalog.yml` with descriptions, targets, dependencies,
risk notices and typed, opt-in run parameters. Unchanged inputs remain inherited.
- Standardizes runnable names as `<category>_<verb>_<purpose>.yml`; retains standard
role entrypoints such as `tasks/main.yml` and `defaults/main.yml`.
- Removes the deprecated `gebhardt` customer playbook and empty `_sample` role.
- Preserves firewall task arguments/order and per-customer policy identities;
extracts reusable role entrypoints and adds preflight validation.
- Adds `aim_debug` for selected diagnostics without enabling `ANSIBLE_DEBUG` or
disabling secret protections. Normal reports remain visible.
- Routes AlmaLinux/Rocky through RedHat-family patching; does not install an
alternate Python interpreter automatically.
- Repairs optional Windows Checkmk YAML generation, configurable variable
precedence, and installed-but-stopped agent service handling.
- Adds known monitoring-script selections and a single managed UniFi configuration
with Vault-backed password references, safe shell quoting and mode-specific URLs.
- Keeps removal of the opposite UniFi check during mode replacement; other cleanup
is a separately approved preview/delete operation.
- Preserves existing controller maintenance, GPO/WinRM provisioning, UI styling,
pagination and inventory/SSH/Vault paths.
- Adds guarded source-only migration and controller-side validation tools.
- Updates the migration guard so operator-approved differences in `scripts/` and `requirements.yml` are backed up and replaced instead of blocking deployment; playbook/role mismatch protection remains strict.
- Adopts the operator-provided unpinned collection baseline and includes `pfsensible.core`.
- Uses staged-role isolation for syntax checks so refactored playbooks cannot accidentally validate against old live roles.
See the migration document for intentionally changed behavior and verification limits.
### Final rc7 Checkmk/script classification hotfix
- Classify `citrix_sessions_customized.ps1`, `veeam_o365_status.ps1`, and `veeam_backup_status.ps1` as AIM-managed Checkmk custom plugins under `$CUSTOM_PLUGINS_PATH$` (`C:\ProgramData\checkmk\agent\plugins`).
- Add `veeam_backup_license_status.ps1` as a VBR-detected Windows local check under `$CUSTOM_LOCAL_PATH$`.
- Stop enabling the Veeam backup status file from `$BUILTIN_PLUGINS_PATH$`; its exact managed rule now targets `$CUSTOM_PLUGINS_PATH$` and follows `want_windows_veeam_backup`.
- Selected custom-plugin deployment removes only known legacy AIM copies from the local directory, plus known historic built-in copies for Veeam O365/backup status.
- Centralize AIM documentation under `scripts/docs/`. The deployment source no longer owns or replaces the installation-root `README.md`.
+247
View File
@@ -0,0 +1,247 @@
# Checkmk deployment and UniFi settings
## 1. Controller maintenance versus host deployment
AIM's top-level Maintenance downloads agent packages and updates the external
monitoring repository directly. No separate Bash helper is installed or required. These operations are separate from customer playbooks.
The package directory and stable filenames remain:
```text
/etc/ansible/roles/checkmk_agent/files/
check_mk_agent.msi
check-mk-agent.deb
check-mk-agent.rpm
```
`checkmk_install_agent.yml` installs the selected platform package, then deploys
checks/configuration and ensures the agent service is running.
`checkmk_update_scripts_config.yml` does not install packages.
`checkmk_cleanup_scripts.yml` previews obsolete managed paths by default.
AIM passes its two non-secret controller paths through environment variables used
at role-default precedence. Inventory overrides still win. Manual execution uses
the documented `/etc/...` defaults unless you configure the role variables.
## Controller-wide Checkmk agent source
AIM stores controller-wide settings in `/etc/ansible/scripts/aim.yml`. The Checkmk agent source is structured rather than stored as a complete URL:
```yaml
maintenance:
checkmk_agent:
protocol: https
base_url: monitoring.domain.de
site_name: monitoring
version: 2.4.0
patchlevel: 46
revision: 1
role_files_dir: /etc/ansible/roles/checkmk_agent/files
```
AIM builds the download root as `https://monitoring.domain.de/monitoring/check_mk/agents` and the package version as `2.4.0p46-1`. `/check_mk/agents` is always appended automatically. These values are edited in the top-level Settings menu and reused by Maintenance without asking again.
## 2. Known external script sources
The repository remains `/etc/checkmk_monitoring_scripts`. AIM performs the Git update itself.
Because this is a managed shadow checkout, AIM uses `git fetch --prune` followed by
`git reset --hard @{u}` rather than invoking an external `git pull` wrapper. Tracked
local drift is discarded after explicit confirmation; untracked files are retained.
File contents are not
included in this bundle and were not supplied for this review; they are not
rewritten, renamed, or executed by the build tests.
| File | Selection |
| --- | --- |
| `Scripts Windows/check-ping.ps1` | Detected domain controller, as before |
| `Scripts Windows/veeam_config_backup_status.ps1` | Detected Veeam VBR, as before |
| `Scripts Windows/veeam_backup_license_status.ps1` | Detected Veeam VBR; local check → `$CUSTOM_LOCAL_PATH$` |
| `Scripts Windows/veeam_o365_status.ps1` | Detected Veeam VBO; custom plugin → `$CUSTOM_PLUGINS_PATH$` |
| `Scripts Windows/citrix_sessions_customized.ps1` | `want_windows_citrix`; custom plugin → `$CUSTOM_PLUGINS_PATH$` |
| `Scripts Windows/veeam_surebackup_status.ps1` | `want_windows_surebackup` |
| `Scripts Windows/windows-backup.ps1` | `want_windows_backup` |
| `Scripts Windows/check-nsp-mailqueue.ps1` | `want_windows_nsp_mailqueue` |
| `Scripts Windows/win_check_cert.ps1` | `want_windows_certificate` |
| `Scripts Windows/veeam_cloud_connect_status.ps1` | `want_windows_veeam_cloud_connect` |
| `Scripts Windows/veeam_backup_status.ps1` | `want_windows_veeam_backup`; custom plugin → `$CUSTOM_PLUGINS_PATH$` |
| `Scripts Linux/check_certificate_directory.sh` | `want_linux_check_certificate` |
| `Scripts Linux/check_unifi-controller/check_unifi-controller.sh` | UniFi network mode |
| `Scripts Linux/check_unifi-controller/check_unifi-os.sh` | UniFi OS mode |
All optional switches default to false. Auto-detected existing behavior is retained.
The repository Veeam backup local check is a separate opt-in from the built-in
agent plugin to avoid silently duplicating it.
`citrix_sessions_customized.ps1`, `veeam_o365_status.ps1`, and `veeam_backup_status.ps1` are custom plugins, not local checks. AIM enables their exact `$CUSTOM_PLUGINS_PATH$` rules before the general
custom-plugin deny rule. When selected, deployment removes the legacy AIM-managed
copy from `C:\ProgramData\checkmk\agent\local` after the plugin copy is in place.
Unknown files remain untouched.
Source-file preflight checks that each selected file exists before deployment.
A filename does not establish the check's internal parameters, privileges or
credentials: configure those according to the maintained script repository. No
new script-specific credential format was guessed from the filenames.
## 3. Single UniFi configuration
```yaml
# host_vars/<fqdn>/main.yml
checkmk_unifi_mode: network # auto, network, os, disabled
checkmk_unifi_username: bf-monitoring
checkmk_unifi_password: "{{ vault_checkmk_unifi_password }}"
```
In the encrypted customer Vault, set:
```yaml
vault_checkmk_unifi_password: "REPLACE_WITH_THE_REAL_MONITORING_PASSWORD"
```
This is the UniFi monitoring password, not the Ansible Windows/SSH login password.
Existing Vaults are not silently modified. Vault creation/consolidation can prepare
an empty optional key; an empty value or `CHANGEME` fails deployment preflight.
Different hosts can override the reference, without exposing the secret in AIM:
```yaml
checkmk_unifi_password: "{{ vault_unifi_site_a_password }}"
```
The AIM run-option field accepts the **variable name** `vault_unifi_site_a_password`,
not a password value. Persist long-lived per-host choices in host vars; run options
are temporary extra variables for the selected execution only.
### Mode and endpoint defaults
| Mode | Check | Default BASEURL |
| --- | --- | --- |
| `network` | `check_unifi-controller.sh` | `https://127.0.0.1:8443` |
| `os` | `check_unifi-os.sh` | `https://127.0.0.1:11443` |
| `auto` | Detected variant; OS wins if both are detected | Variant default |
| `disabled` | Neither selected for deployment | No config replacement |
An explicit `checkmk_unifi_baseurl` overrides the endpoint. In selected network/OS
mode, the role writes one `/etc/check_mk/unifi.cfg`, deploys the matching check,
and removes only the opposite UniFi check. An explicitly disabled mode does not
silently remove existing files during a normal update: review the cleanup action.
### Supplied config schema preserved
```yaml
checkmk_unifi_curl_options: " --insecure --tlsv1.2"
checkmk_unifi_status_provisioning: 1
checkmk_unifi_status_upgrading: 1
checkmk_unifi_status_upgradable: 0
checkmk_unifi_status_heartbeat_missed: 1
checkmk_unifi_status_noautobackup: 0
```
Statuses are 0 OK, 1 WARN, 2 CRIT, 3 UNKNOWN. The template emits USERNAME, PASSWORD,
BASEURL, CURLOPTS and the five STATUS_* settings with shell-safe quoting, validates
shell syntax and sets mode `0600`. Secret-bearing tasks have `no_log: true` and
`diff: false`. The source repository's `unifi.cfg` is no longer copied to the host.
The supplied insecure TLS option remains unchanged; enabling certificate validation
is a separate operational decision, not hidden in this refactor.
## 4. Read current Windows user configuration
Before or after a rollout, `checkmk_read_windows_config.yml` can display the host's current `check_mk.user.yml` without changing it. This is intended for verifying existing `global`, `local`, `mrpe`, plugin, or custom sections before AIM touches the managed `plugins:` section.
```bash
ansible-playbook -i inventories/CUSTOMER/hosts.yml \
playbooks/checkmk_read_windows_config.yml --limit HOST \
--vault-id CUSTOMER@prompt
```
The default path is `C:\ProgramData\checkmk\agent\check_mk.user.yml`. Set `checkmk_windows_user_cfg` only for hosts using a different location. The playbook reports a missing file without creating it and makes no Checkmk changes.
## 5. Cleanup
AIM always starts cleanup in preview mode. It requires explicit deletion enablement
and a second confirmation before running with `checkmk_cleanup_enabled=true`.
Only the known managed filenames are considered; unrelated custom files stay.
The preview is a list of paths considered obsolete by the selected desired state,
not proof each path currently exists. This is distinct from the approved opposite-
UniFi removal during normal mode replacement.
Manual examples:
```bash
ansible-playbook -i inventories/CUSTOMER/hosts.yml \
playbooks/checkmk_cleanup_scripts.yml --limit HOST \
--vault-id CUSTOMER@prompt
# Destructive: enable only after reviewing the intended checks/mode.
ansible-playbook -i inventories/CUSTOMER/hosts.yml \
playbooks/checkmk_cleanup_scripts.yml --limit HOST \
--vault-id CUSTOMER@prompt -e '{"checkmk_cleanup_enabled": true}'
```
## 6. Windows execution policy
AIM now treats `C:\ProgramData\checkmk\agent\check_mk.user.yml` as a shared
operator configuration file, not as an AIM-owned file. On Windows, AIM owns only the
top-level `plugins:` section. The existing `global`, `winperf`, `fileinfo`, `logwatch`,
`local`, `mrpe`, unknown top-level sections, and their comments are preserved. This is
important for older or manually customized Checkmk agents where those sections may
contain host-specific behavior.
AIM keeps an ownership notice on the first line and places another ownership comment
immediately before the `plugins:` section. Only the marked `plugins:` section may be
replaced on later runs. If no `plugins:` section exists, AIM appends one without
rewriting the rest of the file.
The effective AIM plugin timeout remains 120 seconds. Settings live in role defaults so
host/group overrides work. Optional plugin rules are typed list entries such as:
```yaml
checkmk_extra_plugin_patterns:
- pattern: '$CUSTOM_PLUGINS_PATH$\example.ps1'
run: true
async: true
timeout: 90
cache_age: 600
```
AIM deliberately does not manage `local:` or `mrpe:`. Existing local-check execution
policy and MRPE definitions remain exactly under operator/Checkmk control. The same
applies to any other non-plugin user configuration.
The `plugins:` section retains the standard AIM ordering: explicit/custom rules first,
then known built-in rules, custom-plugin catch-alls, the built-in deny rule, and the
final safety deny. Checkmk reads the user configuration after default and Bakery
configuration, so section-scoped ownership avoids unintentionally overriding unrelated
settings.
Actual behavior of the Windows agent and external checks still requires a representative
pilot.
## Structured reports in 3.3.0rc8
Install reports query registry/package database version and current service state after
the existing management roles. Config-update reports list changed sections/check files
and the narrowly approved opposite-UniFi deletion. They do not expose raw config diffs,
guess unknown-file counts or alter unmanaged sections. The reporting-only checkmk_report
role does not install/configure services itself.
The read playbook now limits paths to basename check_mk.user.yml and 512 KiB. Native
terminal readout remains explicit; API reports parse sections and redact recognized
secret/command fields, including MRPE commands. Comments are not parsed YAML data.
See OPERATION_RESULTS.md for exact semantics and VALIDATION.md for acceptance limits.
## Windows AIM-managed file ACLs
AIM normalizes ACLs only on persistent files it manages directly: selected Windows Checkmk scripts in `$CUSTOM_LOCAL_PATH$` / `$CUSTOM_PLUGINS_PATH$` and `check_mk.user.yml`. Unknown files are not touched and directory trees are not rewritten recursively.
For each AIM-managed persistent file, `checkmk_windows_acl` enables parent inheritance and guarantees these well-known SID permissions independent of Windows display language:
| SID | Principal | Rights |
|---|---|---|
| `S-1-5-18` | SYSTEM | FullControl |
| `S-1-5-32-544` | local Administrators | FullControl |
| `S-1-15-2-1` | ALL APPLICATION PACKAGES | ReadAndExecute |
| `S-1-15-2-2` | ALL RESTRICTED APPLICATION PACKAGES | ReadAndExecute |
AIM does not add customer-specific administrator/user ACEs. Parent or explicit ACEs that already exist are not blindly purged; the goal is to guarantee the standard Checkmk-style administrative/application access while preserving operator ACL intent.
+284
View File
@@ -0,0 +1,284 @@
# AIM 3.3.0rc8: safe play/task/host progress
**Service / wire / event version:** 1.0, stable additive 1.x.
**Detail schema:** `play_task_host_v1`. **Canonical Ansible Core:** 2.19.11.
**Qualification:** locally tested with simulated native command/callback fixtures;
real Ansible 2.19.11 and managed-host acceptance is still required for this feature.
## 1. Select the feature without breaking existing clients
Clients receive anonymous progress and aggregate counters by omitting
`progress_mode`. Current Core accepts `summary` (the default) or `detail` in
`RunRequest`. There is no `-vvv`, debug-output forwarding or raw-output option.
Inspect `aimctl capabilities` / `AimService.capabilities()` first:
```json
{
"execution_progress": {
"modes": ["summary", "detail"],
"default": "summary",
"request_field": "progress_mode",
"detail_schema": "play_task_host_v1",
"detail_event_kinds": ["play_started", "play_skipped", "play_stopped", "task_started", "host_result", "task_retry", "task_async_poll", "host_recap"],
"static_source_labels_only": true,
"raw_output": false,
"error_policy": "fixed_diagnostic_hints",
"qualification": "controller_acceptance_required"
}
}
```
This is the relevant capability fragment, not the entire response. Servers without
the advertised detail capability can reject unknown request fields: do not send
`progress_mode` to them. Absence of the capability means use the legacy default or
explain that detailed progress needs a newer core. Administrative enablement and
same-UID runtime readiness remain separate checks.
For a known inventory host, prepare a request with detail selected:
```bash
printf '%s\n' '{"api_version":"1.0","operation":"prepare","request":{"customer":"CUSTOMER","playbook":"debug_test_connection","hosts":["HOST"],"progress_mode":"detail"}}' | aimctl request
```
Substitute real catalog/customer/host identifiers. Input to `aimctl request` is
**one JSON object on one line**, not pretty-printed multiline JSON. Preparation
collects no password and contacts no managed host. Use the same normalized request
(including the mode) and its revision for execution. Changing the mode changes the
reviewed request and requires preparation/review again. An ordinary request never
contains credentials; use the private provider/FD documented in `ADDON_API.md`.
Detailed mode still emits existing `stage`, `progress`, `stats`, and `result` events.
A detail renderer should ignore anonymous `progress` events to avoid duplicate task
lines. An existing summary client can ignore unfamiliar event kinds. Product version
is 3.3.0rc8; API, wire and event version strings stay 1.0.
## 2. Public event envelope and fields
Every event has `event_version`, `run_id`, increasing `sequence`, UTC `timestamp`,
and `kind`. On the wire it is wrapped as `{"type":"event","event":{...}}`.
IDs (`play_id`, `task_id`) are opaque strings scoped to a run. Correlate with IDs,
not names or the last task that happened to arrive. Results can interleave across
hosts under a free strategy. Serial batches can enter the same play again and get
new IDs. Repeated task labels are not unique identities.
| Kind | Additional fields | Meaning |
|---|---|---|
| `play_started` | `play_id`, `label`, `label_redacted` | A play occurrence began; label is unexpanded source text or a fixed placeholder |
| `play_skipped` | `play_id`, `reason: "no_hosts_matched"` | Native no-matching-host callback; not an authentication failure |
| `play_stopped` | `play_id`, `reason: "no_hosts_remaining"` | Native no-hosts-remaining callback; not equivalent to a harmless no-match skip |
| `task_started` | `play_id`, `task_id`, `label`, `label_redacted`, `handler` | A task/handler started; no arguments, module name, role path, source path or templated name |
| `host_result` | `play_id`, `task_id`, `host`, `host_redacted`, `status`, `changed`, `ignored`, `details_redacted`, `error` | One aggregate host/task outcome |
| `task_retry` | `play_id`, `task_id`, `host`, `host_redacted`, `attempt`, `details_redacted` | Native task retry, not AIM replay of the playbook |
| `task_async_poll` | Same fields as `task_retry` | Async poll notification; no job ID or result payload |
| `host_recap` | `host`, `host_redacted`, `counts` | Final counts for one logical inventory host |
`host_result.status` is `ok`, `changed`, `skipped`, `failed`, or `unreachable`.
`changed` is a boolean and can also be true on a failed task. `ignored` reflects
native ignore-error/ignore-unreachable handling; a failed event is not by itself
proof that the whole run failed. Rescue and ignored counts remain in final stats.
`attempt` is an integer or null; it is always null when details are withheld.
`host` is a reviewed logical inventory name, never a resolved connection address,
delegation target, loop label or variable. Out-of-scope/unsafe names are represented
as null with `host_redacted: true`. This can occur for trusted delegated/dynamic
inventory work: host limits are an operational selection, not a security sandbox.
A host name containing a supplied credential is also withheld.
Every `counts` mapping has nonnegative integers for `ok`, `changed`, `failures`,
`unreachable`, `skipped`, `rescued`, `ignored`. Host recaps are followed by the
existing aggregate `stats` and final run result. Do not count progress lines to
calculate a recap, assume one task header per host, or infer success from a zero
process exit without the authoritative final response/result. A failure followed
by rescue can still result in a successful native run.
There is no invented `play_completed` event: use the actual following play/final
stats boundary. Loops produce aggregate host/task outcomes, not item values/events.
Async poll and retry are optional native notifications; fire-and-forget async work
is not thereby certified complete. The console's exact visual layout is not the API.
Example host outcome (IDs/timestamp illustrative):
```json
{
"event_version": "1.0",
"run_id": "example-run",
"sequence": 12,
"timestamp": "2026-09-19T12:00:00+00:00",
"kind": "host_result",
"play_id": "p2",
"task_id": "t1",
"host": "host01.example.test",
"host_redacted": false,
"status": "unreachable",
"changed": false,
"ignored": false,
"details_redacted": false,
"error": {
"code": "connection_refused",
"message": "The connection was refused. Check the target listener, port and firewall.",
"classification": "diagnostic_hint"
}
}
```
## 3. Failure information without raw error text
A host failure carries `error: {code, message, classification}`. Success/skip events
carry `error: null`. Messages are fixed strings owned by core. They never contain
the original result's `msg`, module arguments, exception, stdout or stderr.
| Code | Interpretation |
|---|---|
| `connection_refused` | Native unreachable message matched a connection-refusal signature; check listener, port and firewall |
| `connection_timeout` | Native unreachable message matched a connection timeout |
| `name_resolution_failed` | Native unreachable message matched a DNS/name-resolution error |
| `tls_verification_failed` | Native unreachable message matched certificate verification failure |
| `authentication_failed` | Native unreachable message matched a recognized authentication rejection |
| `permission_denied` | Failed task message matched an access/permission denial |
| `host_unreachable` | No narrower supported hint; transport/authentication details remain withheld |
| `task_failed` | No narrower supported hint; arbitrary module error text remains withheld |
| `details_withheld` | Sensitive/no_log result; no diagnostic classification is exposed |
Hints are derived from bounded native message signatures, **not a definitive root
cause or proof that a supplied password was used**. Localized/unrecognized messages
can produce generic errors. Never drive automatic credential retries, disable TLS
validation, open firewall rules or replay operations from these hints. The existing
RunResult error/remote-work flag remains authoritative for lifecycle decisions.
Preflight errors still use the existing fixed service error codes; this is not a
raw syntax-error/traceback channel. More detail may require an authorized operator's
trusted terminal diagnostics, handled as potentially sensitive data.
## 4. Safety contract and unavoidable trust boundary
Core reads the original parsed play/task `name` field, not Ansible's templated
`get_name()` or rendered task fields. Missing, templated, overlong, control-bearing,
URL-bearing or credential-assignment-like labels are replaced by fixed labels.
Known supplied credential values matching a label are also withheld at the public
boundary. Labels have at most 200 characters / 800 UTF-8 bytes; public host names
at most 255 characters. Longer valid inventory identifiers can execute, but their
name is withheld in the detail stream.
Known true or potentially true `no_log` on a task, block, role or play suppresses
its label. For dynamic no_log expressions core does not attempt to render the
expression. Results marked `no_log`, censored, or containing hidden loop results
retain only status/changed/ignored, allowed host and IDs. Their failures use
`details_withheld`. Retry attempts are hidden for these results.
**Static source labels and inventory names must themselves be non-secret.** No
filter can identify every secret literally hard-coded in a name. A runtime-only
sensitivity flag also cannot retroactively retract a static header already sent.
This is a trusted controller/source contract, not a general secret-scanning engine
or malicious-playbook sandbox. Add-on authors must never put credentials or
variable values in labels. Dynamic names deliberately lose their expansion.
Nothing here authorizes raw `debug` values, Checkmk configuration contents,
registered variables, `invocation`, environment/command strings, loop items,
exceptions, custom stats, module stdout/stderr, or `-v/-vv/-vvv/-vvvv` output.
`raw_task_output` remains unsupported. Known password masking is defense in depth,
not permission to pass arbitrary output through a redactor.
Core validates the private callback schema, IDs, sequences and counts before
building public events. Detail collection is bounded to 200,000 private frames
(including legacy counters/handshakes), 50,000 play occurrences and 50,000 task IDs.
These are safety bounds, not an estimated progress denominator. Malformed, truncated,
out-of-order or incomplete streams cannot produce a successful run. Typical core
errors are `invalid_event_stream`, `event_limit`, `event_bridge_unavailable`, and
`event_bridge_incomplete`. A native failure may end before a complete recap: retain
received events and show failure/unknown, never fabricate the missing recap.
The event sink must remain responsive or enqueue into a bounded queue; it must not
block on slow browsers. Core still uses the established process deadlines and
cancellation. No output detail mode changes account permissions, credential
precedence, execution enablement, or automatic retry policy.
## 5. Example plain-text renderer for an independent client
This is an illustrative client function, not a new core CLI command. Feed it
validated public event objects from your client after version/capability checks.
Use text nodes/HTML escaping in a browser, not `innerHTML`; escape rich terminal
markup as well. Treat labels as text, not as commands or markup. Persist only the
minimal events your deployment policy needs and restrict job-log visibility.
```python
class ProgressText:
def __init__(self):
self.tasks = {}
self.active_task = None
self.recap_started = False
def __call__(self, event):
kind = event.get('kind')
if kind == 'play_started':
print('\nPLAY [' + event['label'] + ']', flush=True)
self.active_task = None
elif kind == 'play_skipped':
print('skipping: no hosts matched', flush=True)
elif kind == 'play_stopped':
print('stopped: no hosts remaining', flush=True)
elif kind == 'task_started':
self.tasks[event['task_id']] = event['label']
self.active_task = event['task_id']
prefix = 'HANDLER' if event['handler'] else 'TASK'
print('\n' + prefix + ' [' + event['label'] + ']', flush=True)
elif kind == 'host_result':
task_id = event['task_id']
# Results can interleave: repeat the correct header, not the last name.
if self.active_task != task_id:
print('\nTASK [' + self.tasks.get(task_id, 'Task') + ']', flush=True)
self.active_task = task_id
host = event['host'] or 'host withheld'
suffix = ' (ignored)' if event['ignored'] else ''
error = event.get('error')
if error:
suffix += ' => ' + error['message']
print(event['status'] + ': [' + host + ']' + suffix, flush=True)
elif kind in ('task_retry', 'task_async_poll'):
print(kind + ': [' + (event['host'] or 'host withheld') + ']', flush=True)
elif kind == 'host_recap':
if not self.recap_started:
print('\nPLAY RECAP', flush=True)
self.recap_started = True
counts = ' '.join(k + '=' + str(v) for k, v in event['counts'].items())
print((event['host'] or 'host withheld') + ' : ' + counts, flush=True)
# Ignore legacy anonymous progress to avoid duplicate lines in detail mode.
# The caller handles stage/stats/result and the authoritative final response.
```
A GUI timeline should key results by `(run_id, play_id, task_id, host)` where the
host is visible. It must not merge all withheld-host records as one known machine.
Keep the original event sequence for audit; a display regrouping is presentation
only. The built-in `aim` terminal still uses its existing native Ansible output,
not this example renderer. No add-on source was changed for this feature.
## 6. Native contracts consulted (not proof of native execution)
Implementation targets Ansible Core 2.19.11 callback entry points, original source
mappings and CallbackTaskResult public properties. Any dependence on these internals
stays in the core-owned adapter, not in an add-on.
- https://docs.ansible.com/projects/ansible-core/2.19/plugins/callback.html
- https://docs.ansible.com/projects/ansible-core/2.19/reference_appendices/logging.html
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/plugins/callback/default.py
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/executor/task_result.py
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/executor/playbook_executor.py
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/playbook/base.py
Follow `SANITY.md` for exact-runtime/controller acceptance before treating the
new timeline as qualified in an installed service or add-on deployment.
## Final target outcomes
Do not reconstruct final requested-target success from this live detail stream.
`RunResult.target_summary` and `RunResult.targets` now provide Core-owned final
accounting under `target_outcome_summary_v1` in both summary and detail modes. The
detail stream remains for live presentation and diagnostics; the final target
summary is authoritative for per-request host outcomes. See `TARGET_OUTCOMES.md`.
## Purposeful data is separate
The 3.3.0rc8 `operation_result` channel does not loosen this host/task event contract.
Only explicitly catalogued, typed and validated set_stats data is returned in the
final result; see OPERATION_RESULTS.md. Arbitrary debug output and configuration
bodies remain excluded from progress events. The Checkmk reader now publishes parsed
redacted sections through its own declared report, not raw debug text.
+176
View File
@@ -0,0 +1,176 @@
# Controller staging and hardened executors
**AIM 3.3.0rc8 / API 1.0. Profile: native_defaults_preflight_v2.**
This is a generic executor deployment requirement, not an add-on-specific policy.
## Required default
Core separately checks the process/config HOME used by controller local_tmp and the
passwd/NSS account home used by default POSIX local/delegated-local tilde expansion.
When these differ, both private `.ansible/tmp` locations may need narrow writable
exceptions inside the service sandbox. A successful ordinary shell check is not proof
of in-service access. Configured absolute local_tmp paths are checked separately.
Keep owner-only 0700 directories owned by the executor (an inherited setgid bit is
accepted without repair). Do not grant group/other access, writable whole-home trees
or disable ProtectSystem/ProtectHome. Core does not fix ownership, ACLs, units or mounts.
The profile creates/writes/reads/removes its own probe files and may create missing
private staging directories. It does not purge existing content. Both paths must be
traversable without symlink components. Inventory-specific connection/become/remote_tmp
overrides remain separate qualification work; this is not a complete Ansible variable
resolver. See VALIDATION.md for the scope of historical controller evidence.
## Always-on Core preflight
Core now checks staging during `readiness` and `execute` before native-runtime
inspection and before consuming any credentials. It checks again immediately
before playbook launch. A local failure returns a specific readiness error rather
than launching remote work and discovering it at a delegated controller task.
Neither `capabilities`, metadata discovery nor `prepare` performs the probe.
The protected local operation is also available without a customer or Ansible:
```bash
aimctl staging-check
printf '%s\n' '{"api_version":"1.0","operation":"staging_check"}' | aimctl request
```
Python: `AimService().staging_check()`. It requires the configured active AIM group
and ordinary filesystem access, but not `addons.execution_enabled`. It does not
accept credentials, path overrides, a customer, target or arbitrary command.
The probe uses the selected AIM Ansible configuration and filtered service
environment. It runs as an isolated Python child under the **same UID, groups and
namespace**, bounded to ten seconds. It creates missing directory components with
0700 access, then creates a unique scratch child, writes/flushes/reads a fixed
non-secret file and removes that file and scratch child. Newly created base
directories remain, as native Ansible also needs them. No existing contents are
listed, read, deleted or chmod/chowned. On abrupt process/storage failure a random
`.aim-staging-check-*` child containing only the fixed non-secret marker may remain;
Core never cleans up another process's children recursively.
Success returns `api_version`, `profile: native_defaults_preflight_v2`, `ready`,
`uid`, `gids`, `directories`, `process_home`, `account_home`, `home_paths_differ`,
`remote_connected: false`, `credentials_read: false`, `remote_paths_overridden: false`,
and coverage warnings. Each directory has a
`purpose`, `source`, resolved `path`, `writable`, `private`, and
`created_directories` flag. If two purposes use one path it is probed once but both
purposes are reported. Paths are operator metadata; do not place credentials in
path names. Readiness nests this report under `runtime.controller_staging`.
Fixed errors have `stage: readiness`, `retryable: false`, and no required secrets:
| Code | Action |
|---|---|
| `controller_staging_unavailable` | Review disk/mount/access restrictions and allow the exact private staging directory in the executor sandbox |
| `controller_staging_unsafe` | Review ownership, owner-only access, directory type or symlinks; no automatic repair is attempted |
| `controller_staging_config_unsupported` | Use supported literal configuration paths; do not assume unresolved expressions were checked |
| `controller_staging_timeout` | Inspect local storage/sandbox before a fresh attempt |
| `controller_staging_probe_failed` | Inspect interpreter/sandbox availability; no raw child errors are forwarded |
An `execute` failure at this gate has `remote_work_may_have_started: false`, null
native `exit_code`, and no fabricated unreachable counts. A bad staging path is
not a password rejection. A second check after credential/syntax preflight can
fail without replaying or launching the playbook. No automatic retry is added.
## What is and is not resolved
For the native controller temporary directory, the probe honors literal
`[defaults] local_tmp` in the selected `ANSIBLE_CONFIG`; otherwise it uses
`[defaults] home` plus `/tmp`, or the native `~/.ansible/tmp` default. Supported
paths are absolute, `~/...`, `$HOME/...`, `${HOME}/...` and literal configuration-
relative paths. Matching outer quotes are removed. Relative controller paths are
relative to the configuration file. Jinja, other environment expansions, `~user`,
parent traversal, wildcard/control-bearing and root/shared-system-root paths are
rejected instead of guessed. The configuration is never rewritten.
The probe also checks the default POSIX `connection: local` module-staging path when `[defaults] remote_tmp` is absent or explicitly has that default. Ansible local connection expands `~user` from the execution account database, so Core resolves this second purpose from the account passwd/NSS home rather than process `$HOME`. When those homes differ, both paths are reported and probed separately. If a different global `remote_tmp` exists, Core reports that its
controller-delegation use is **not probed**. It must not create a possibly remote
path on the controller. Inventory-specific `ansible_remote_tmp`, delegated
connection/shell/interpreter/escalation overrides, caches, async directories and
module execution permissions are not fully resolved by this credential-free
probe. Passing it is not complete playbook or sandbox qualification.
No environment, Ansible command, inventory or remote path is changed. In
particular, Core does not set global `ANSIBLE_REMOTE_TEMP`, force local connection,
enable local become, redirect HOME or inherit arbitrary environment variables.
The established service filter still drops inherited `ANSIBLE_LOCAL_TEMP` and
`ANSIBLE_REMOTE_TEMP`; a unit environment assignment alone is not configuration
of this execution profile. Use the operator-owned selected Ansible configuration
for supported local_tmp changes and explicitly review local delegation separately.
## Standard systemd deployment recipe for add-on owners
New add-on installers that provision an AIM executor should apply the equivalent
of this recipe by default. Core publishes the requirement; it does **not** discover,
edit, reload or restart independently managed service units. Do not install it on
a browser/frontend process that does not execute AIM. Never give a frontend private
key access just to satisfy an executor check.
1. Resolve the actual executor account's home with NSS/getpwnam. Do not assume that
a system unit's `%h` denotes its `User=` account: use the resolved absolute path.
2. Before activating the sandbox, provision the private staging location as that
account, or run the following local probe from its ordinary administrative context:
```bash
# Replace EXECUTOR_USER and the AIM launcher/config paths for the deployment.
sudo -u EXECUTOR_USER -H /usr/local/bin/aimctl staging-check
```
This can create the missing private directories if their parents permit it. It
will not fix pre-existing ownership or permissions. Outside-unit success only
checks provisioning; it does not qualify the service sandbox.
3. Preserve existing unit settings and add an operator-owned drop-in like:
```ini
[Service]
# Required default for AIM controller-local and implicit-localhost staging.
# Substitute the real absolute executor home; do not copy the placeholder.
ReadWritePaths=/home/EXECUTOR_USER/.ansible/tmp
# If process HOME is deliberately elsewhere and staging-check reports a second
# controller_local_tmp path, permit that exact private path too unless it is
# already covered by an existing writable StateDirectory/ReadWritePaths entry.
# Use the actual core launcher, and --config before staging-check if nondefault.
ExecStartPre=/usr/local/bin/aimctl staging-check
```
These nonempty assignments append to existing lists. Do **not** insert blank
`ReadWritePaths=` or `ExecStartPre=` assignments that reset existing protections or
startup checks. Do not prepend `-` to the staging path/check: a missing/unusable
required directory must fail visibly. Provision it before systemd constructs the
write exception. Keep `ProtectHome=read-only`, `ProtectSystem=strict`,
`PrivateTmp=yes`, `NoNewPrivileges=yes`, and other existing hardening.
`ExecStartPre` runs with the service's identity/sandbox; allow the real AIM Python
and probe subprocesses in any syscall/exec restrictions. No escalation is needed.
If a non-systemd orchestrator is used, provide the equivalent writable mount and
run the same check in the actual worker context.
4. Quiesce jobs, reload the unit and restart it in an approved window. Never restart
a worker in the middle of a deployment. Check its startup result and mount view,
then run one explicitly reviewed controller-delegated operation. Do not globally
remount the home, enable root execution, relax key permissions or chown the tree.
## Existing corrected deployments
Keep the successful narrow write exception. Installing this Core release does
not remove it, duplicate it, or write an add-on unit. The new preflight becomes
automatic on the next API readiness/execution call. Add the optional startup check
through your service's owner so a bad deployment fails at worker startup as well.
Do not remove the exception to test a failure on an active controller. Use a
separate fixture/unit for negative acceptance.
Rollback Core source normally; service overrides are independently managed and
are not changed by Core rollback. Retaining the narrow exception is needed by
older Core releases too. No rollback undoes a completed Checkmk installation.
## Primary references
Consulted native contracts; not evidence of live testing in the release builder:
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/config/base.yml
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/plugins/doc_fragments/shell_common.py
- https://raw.githubusercontent.com/ansible/ansible/v2.19.11/lib/ansible/config/manager.py
- https://raw.githubusercontent.com/systemd/systemd/v257/man/systemd.exec.xml
+484
View File
@@ -0,0 +1,484 @@
# AIM Fresh Installation
**Applies to:** AIM 3.3.0rc8\
**Canonical Ansible Core:** 2.19.11\
**Default controller root:** `/etc/ansible`\
**Supported AIM Python:** Python 3.11+
This guide is for a new AIM controller with no existing AIM source tree.
It installs the complete source bundle, AIM's Python runtime
dependencies, the separate canonical Ansible runtime, and the required
Ansible collections.
The AIM deployer deliberately does **not** install operating-system
packages, Python dependencies, Ansible, collections, accounts, groups,
services, or add-on executors. Provision those explicitly, then use the
bundled deployer in `install` mode.
AIM remains usable as a standalone product through the built-in `aim`
terminal. Add-on execution is optional and disabled by default.
## 1. Obtain and verify the release
Transfer both files through a trusted channel:
``` text
AIM-Ansible-3.3.0rc8.zip
AIM-Ansible-3.3.0rc8.zip.sha256
```
The SHA-256 sidecar verifies file integrity; it is not a publisher
signature.
Stage the release outside `/etc/ansible`:
``` bash
cd /var/tmp
sha256sum -c AIM-Ansible-3.3.0rc8.zip.sha256
unzip AIM-Ansible-3.3.0rc8.zip
cd aim-core-3.3.0rc8
```
Do not extract the release directly over `/etc/ansible`.
## 2. Provision the controller prerequisites
The controller needs:
- Linux
- Python 3.11 or newer for AIM
- a separate Ansible runtime containing exactly
`ansible-core==2.19.11`
- the collections declared by `requirements.yml`
- an operator-approved execution/service account and group policy
- network and remote authentication prerequisites appropriate to the
managed systems
For the default configuration shipped in this release, review
`scripts/aim.yml` before installation. In particular, these values are
environment-specific and must not be accepted blindly:
``` yaml
root_dir: /etc/ansible
service_user: svc_bf-ansible
required_group: srv_debsansible01_admins@bitformer.lan
```
Set them to the intended controller values after installation, or
prepare a reviewed configuration before first operational use.
## 3. Create the AIM Python environment
Keep AIM's Python environment separate from the canonical Ansible
runtime and from add-on environments.
Example:
``` bash
sudo python3.11 -m venv /opt/aim/venv
sudo /opt/aim/venv/bin/python -m pip install --upgrade pip
sudo /opt/aim/venv/bin/python -m pip install \
'ruamel.yaml>=0.18,<0.19' \
'rich>=13,<15'
```
Verify it:
``` bash
/opt/aim/venv/bin/python --version
/opt/aim/venv/bin/python -c 'import ruamel.yaml, rich; print("AIM dependencies OK")'
```
The deployer uses this interpreter for the `aim` and `aimctl` launchers.
It does not install or upgrade these dependencies itself.
## 4. Create the canonical Ansible runtime
AIM 3.3.0rc8 is qualified against **ansible-core 2.19.11**.
One clean layout is a separate virtual environment:
``` bash
sudo python3.11 -m venv /opt/ansible/venv
sudo /opt/ansible/venv/bin/python -m pip install --upgrade pip
sudo /opt/ansible/venv/bin/python -m pip install 'ansible-core==2.19.11'
```
Verify the exact version:
``` bash
/opt/ansible/venv/bin/ansible-playbook --version
```
The first version line must report Ansible Core 2.19.11.
Do not silently substitute a newer Ansible Core release.
## 5. Install the required Ansible collections
The source bundle declares:
``` yaml
collections:
- ansible.netcommon
- name: ansible.windows
version: ">=3.8.0,<4.0.0"
- ansible.posix
- community.windows
- community.general
- sophos.sophos_firewall
- pfsensible.core
```
Install them with the canonical Ansible runtime:
``` bash
sudo /opt/ansible/venv/bin/ansible-galaxy collection install \
-r requirements.yml
```
AIM 3.3.0rc8 requires `ansible.windows>=3.8.0,<4.0.0` for native reboot-state discovery. Other collection versions remain operator-approved and must be compatible with Ansible Core 2.19.11. AIM does not download or upgrade collections during source deployment.
Verify collection discovery:
``` bash
/opt/ansible/venv/bin/ansible-galaxy collection list
```
If collections are installed into a nonstandard path, configure
`runtime.ansible_collections_path` in `scripts/aim.yml` after
installation.
## 6. Preview the fresh installation
For a fresh controller use `install`, not `update`.
The default installation root is `/etc/ansible`, and the default
launcher directory for a fresh installation is `/usr/local/bin`.
Run a dry-run first:
``` bash
sudo python3 deploy/deploy.py install \
--aim-python /opt/aim/venv/bin/python \
--dry-run
```
Review the complete plan. It should show source installation under
`/etc/ansible` and creation of the `aim` and `aimctl` launchers.
If using a nondefault controller root, provide both an explicit target
and a dedicated command directory:
``` bash
sudo python3 deploy/deploy.py install \
--target /srv/aim-controller \
--bin-dir /usr/local/libexec/aim-controller \
--aim-python /opt/aim/venv/bin/python \
--dry-run
```
Do not use a production launcher directory for a second development/test
installation.
## 7. Apply the installation
For a genuinely fresh tree there should be no active AIM jobs, but the
deployer still requires explicit acknowledgement that source
writers/runners are quiesced:
``` bash
sudo python3 deploy/deploy.py install \
--aim-python /opt/aim/venv/bin/python \
--apply \
--quiesced
```
Then refresh shell command discovery:
``` bash
hash -r
command -v aim
command -v aimctl
aim --version
aimctl --version
```
Expected product version:
``` text
3.3.0rc8
```
The deployer verifies the installed launchers and source before
reporting success.
## 8. Configure AIM
Edit the installed controller configuration:
``` bash
sudoedit /etc/ansible/scripts/aim.yml
```
At minimum review:
``` yaml
root_dir: /etc/ansible
service_user: <approved execution account>
required_group: <approved local/NSS group>
runtime:
ansible_playbook: /opt/ansible/venv/bin/ansible-playbook
ansible_collections_path: ""
private_key_owner: ""
addons:
execution_enabled: false
```
Keep `addons.execution_enabled: false` until add-on/service execution
has been deliberately provisioned and accepted.
`runtime.private_key_owner` affects newly generated keys only. It does
not switch the execution UID, migrate existing keys, or grant filesystem
access.
Checkmk-specific controller settings under `maintenance.checkmk_agent`
should be configured only when that functionality is required.
## 9. Establish inventories and customer data
A fresh source installation does not create production inventories,
Vault data, SSH keys, customer credentials, or customer-local
configuration.
The standard controller layout is:
``` text
/etc/ansible/
├── scripts/
│ ├── aim.yml
│ └── ...
├── playbooks/
├── roles/
├── inventories/
└── requirements.yml
```
Customer inventories, Vault files, keys, and `.aim.yml` files are
runtime/operator data and are intentionally absent from the release
archive.
Create or migrate them through the approved AIM/operator workflow. Never
copy credentials into the source release or documentation.
## 10. Verify Core discovery
Run:
``` bash
aimctl capabilities
```
Confirm that the product/API information is correct and that expected
capabilities are advertised.
Then inspect the installed documentation:
``` text
/etc/ansible/scripts/docs/README.md
/etc/ansible/scripts/docs/SANITY.md
/etc/ansible/scripts/docs/VALIDATION.md
```
`VALIDATION.md` records release-development evidence. It is not proof
that this newly installed controller has passed managed-host acceptance.
## 11. Verify controller staging
Before service/add-on execution, check controller-local Ansible staging:
``` bash
aimctl staging-check
```
For a normal interactive/native installation, AIM checks the relevant
controller staging paths.
For a hardened service executor, run this check **inside the actual
service sandbox under the real execution identity**. A successful
`sudo -u` shell test does not reproduce systemd mount restrictions.
With native defaults, both of these may matter:
``` text
<process HOME>/.ansible/tmp
<execution account passwd/NSS home>/.ansible/tmp
```
If the service uses `ProtectHome=read-only`, retain that hardening and
add only the narrow writable exception required for the private staging
directory. Do not make the whole home writable, recursively change
ownership, or disable service hardening.
See:
``` text
scripts/docs/EXECUTOR_STAGING.md
```
## 12. Test the built-in terminal first
AIM is an independent product. Validate the native terminal before
introducing an add-on:
``` bash
aim
```
Use a controlled test customer/host and follow `scripts/docs/SANITY.md`.
Confirm inventory discovery, credentials/Vault handling, target
selection, and an approved low-risk playbook before broader production
use.
## 13. Optional add-on/service execution
Add-on execution is **disabled by default**.
Only after the controller and executor identity have been accepted
should the operator deliberately enable it:
``` yaml
addons:
execution_enabled: true
runtime:
ansible_playbook: /opt/ansible/venv/bin/ansible-playbook
```
The public integration boundary is:
``` text
aim.services.v1
aimctl
```
Do not have an add-on:
- import `aim.ui`
- parse AIM terminal output
- patch Core
- rewrite native Ansible commands
- read/export private keys
- invent its own inventory hierarchy or target outcomes
- reconstruct structured operation results from debug/stdout
The current add-on contracts are documented in:
``` text
scripts/docs/ADDON_API.md
scripts/docs/ADDON_SUPPORT.md
scripts/docs/OPERATION_RESULTS.md
ADDON_AGENTS.md
```
The v1 execution profile is same-UID. Cross-UID execution is not
provided by this release.
## 14. Run the controller acceptance checklist
Use the current checklist:
``` text
scripts/docs/SANITY.md
```
For 3.3.0rc8, acceptance should include the structured operation-result
path in addition to ordinary execution.
A useful first reporting test is host-role detection. Confirm that the
final result exposes the declared structured capability booleans without
requiring the client to parse task output.
Also exercise, as applicable:
- inventory hierarchy
- target outcome summaries
- controller staging
- summary and detail progress
- disk-usage reporting
- service recovery reporting
- Checkmk reporting
- Linux/Windows patch reporting
- successful and mixed-result multi-host runs
Do not promote a release candidate based only on source parsing or local
simulated callback tests.
## 15. Recovery
Fresh installation still creates protected deployment recovery data. The
deployer prints its exact recovery directory under the default:
``` text
/var/backups/aim-core/
```
If installation validation fails, use the same release deployer and
printed recovery directory:
``` bash
sudo python3 deploy/deploy.py rollback \
--from-backup /var/backups/aim-core/RECOVERY-DIRECTORY \
--dry-run
sudo python3 deploy/deploy.py rollback \
--from-backup /var/backups/aim-core/RECOVERY-DIRECTORY \
--apply \
--quiesced
```
A source rollback does not reverse remote Ansible work, dependency
installation, account/group changes, or add-on state.
## Fresh-install completion checklist
A fresh controller is not complete merely because `aim --version` works.
Before production use confirm:
- release ZIP checksum verified
- AIM Python 3.11+ environment provisioned
- AIM Python dependencies installed
- canonical Ansible Core 2.19.11 provisioned separately
- required Ansible collections installed and discoverable
- full source installed with `deploy.py install`
- `aim` and `aimctl` resolve to the intended installation
- `scripts/aim.yml` reviewed for this controller
- approved execution identity/group configured
- inventory/Vault/key data established separately
- `aimctl capabilities` succeeds
- `aimctl staging-check` succeeds in every execution context
- built-in terminal workflow accepted
- managed-host sanity tests completed
- add-on execution remains disabled unless separately provisioned and
accepted
- recovery location recorded and protected
## Related documentation
- existing root `README.md` --- operator-owned and preserved; AIM does not overwrite it
- `deploy/README.md` --- deployment and rollback mechanics
- `scripts/docs/README.md` --- current documentation index
- `scripts/docs/SANITY.md` --- controller acceptance checklist
- `scripts/docs/VALIDATION.md` --- release-development evidence and
limits
- `scripts/docs/EXECUTOR_STAGING.md` --- hardened executor staging
requirements
- `scripts/docs/ADDON_API.md` --- public add-on API
- `scripts/docs/ADDON_SUPPORT.md` --- current add-on support matrix
- `scripts/docs/OPERATION_RESULTS.md` --- structured operation-result
contract
- `scripts/docs/PLAYBOOKS.md` --- catalog/playbook behavior
- `scripts/docs/CHECKMK.md` --- Checkmk-specific behavior
+69
View File
@@ -0,0 +1,69 @@
# AIM inventory hierarchy discovery
**AIM 3.3.0rc8 / service API 1.0.** `inventory_hierarchy_v1` is a read-only,
additive discovery capability. It exists so interfaces can present AIM's actual
inventory group/subgroup structure without duplicating Ansible-inventory parsing.
## Public operation
Python:
```python
result = AimService().inventory_hierarchy('CUSTOMER')
```
Machine interface:
```bash
printf '%s\n' '{"api_version":"1.0","operation":"inventory_hierarchy","customer":"CUSTOMER"}' | aimctl request
```
Query `capabilities.inventory_hierarchy` before relying on this operation.
The result has this shape:
```json
{
"api_version": "1.0",
"schema": "inventory_hierarchy_v1",
"customer": "CUSTOMER",
"hosts": ["direct-customer-host.example"],
"groups": [
{
"name": "windows",
"path": ["windows"],
"hosts": ["win01.example"],
"children": [
{
"name": "servers",
"path": ["windows", "servers"],
"hosts": ["win02.example"],
"children": []
}
]
}
]
}
```
`hosts` on each node means **direct membership in that node**, not recursively
expanded membership. A client can render descendants without guessing whether a
host came from the parent or a subgroup. Group nesting is preserved to a bounded
64 levels / 10,000 group nodes; malformed mappings, aliases/cycles or larger
structures fail as invalid source rather than being silently flattened.
## Safety and execution boundary
Only customer/group names, group paths and inventory host names are returned.
Variables, addresses, credentials, Vault values, connection settings and group
vars are not part of this tree. `list_hosts` remains the flat host/address/platform
metadata operation.
The hierarchy is presentation/discovery metadata only. It does **not** add group
patterns to `RunRequest`, change `--limit`, authorize targets, or skip preparation
and review. Execution still accepts explicit host names and Core revalidates them
against current inventory/catalog compatibility.
Interfaces should not persist a competing hierarchy or derive execution authority
from a previously fetched tree. Fetch current hierarchy when presenting inventory,
then prepare the final explicit host request through Core.
+276
View File
@@ -0,0 +1,276 @@
# Purposeful operation results
**AIM 3.3.0rc8. Publisher `aim_output_v1`; public `aim_operation_result_v1`.**
This is independent of `play_task_host_v1` progress and `target_outcome_summary_v1`.
The service, wire and event versions remain 1.0. It is not a raw-output/debug option.
## Capability and catalog negotiation
`capabilities.operation_results` advertises publisher/result protocols, scopes, modes,
limits and qualification. Each catalog operation has `result: null` or a declaration:
```yaml
result:
protocol: aim_output_v1
schema: host_capabilities_v1
scope: per_host
required: true
sensitivity: safe
max_bytes_per_host: 1048576
schema_file: host_capabilities_v1.yml
```
`schema_file` is a basename under the catalog's `schemas/` directory. Core resolves it
into `data_schema` for public `list_playbooks` and `PreparedRun.result_contract` output.
Clients never pass a schema file or payload in an execute request. Schema/source bytes
participate in the review revision. A changed schema requires preparation/review again.
No hard-coded playbook dispatch exists in Core.
## Publisher convention
```yaml
- name: AIM | Publish operation result
ansible.builtin.set_stats:
per_host: true
aggregate: false
data:
aim_output:
protocol: aim_output_v1
schema: host_capabilities_v1
data:
is_dc: '{{ is_dc | default(false) | bool }}'
is_dhcp_server: '{{ is_dhcp_server | default(false) | bool }}'
is_hyperv_host: '{{ is_hyperv_host | default(false) | bool }}'
has_veeam_vbr: '{{ has_veeam_vbr | default(false) | bool }}'
has_veeam_vbo: '{{ has_veeam_vbo | default(false) | bool }}'
has_veeam_em: '{{ has_veeam_em | default(false) | bool }}'
is_unifi_controller: '{{ is_unifi_controller | default(false) | bool }}'
is_unifi_os_server: '{{ is_unifi_os_server | default(false) | bool }}'
```
Use exactly one publication per requested host/run. Do not loop publishers, publish
preliminary values under the same key, or aggregate dictionaries accidentally. For a
global declaration use `per_host: false` and exactly one global publication (beware
run_once with serial batches). The bundled schemas are all per-host.
The callback observes only an actual set_stats result with a matching declaration,
false aggregation and matching scope. It records sensitivity provenance and reconciles
that value with final custom statistics. An unrelated custom stat, a debug task named
"Publish", or a forged final statistic is not an output source. `no_log` publishers
are withheld; custom stats alone cannot erase their provenance. Report from a reviewed
normalizer with selected non-secret fields, not arbitrary registered result dictionaries.
Publishing set_stats is controller-side and adds no remote package/module dependency.
Native console rendering remains unchanged; use the existing debug summaries for CLI
operators. External clients obtain reports only at finalization, not through debug text.
## Final public object
Every RunResult adds `operation_result`, null for an undeclared operation. Otherwise:
```json
{
"protocol": "aim_operation_result_v1",
"schema": "host_capabilities_v1",
"scope": "per_host",
"required": true,
"complete": true,
"check_mode": false,
"hosts": {
"host01.example": {
"schema": "host_capabilities_v1",
"status": "available",
"data": {"is_dc": false, "is_dhcp_server": false, "is_hyperv_host": false,
"has_veeam_vbr": false, "has_veeam_vbo": false, "has_veeam_em": false,
"is_unifi_controller": false, "is_unifi_os_server": false},
"error": null
}
},
"global": null
}
```
For global scope, `hosts` is empty and `global` holds the same entry shape. Schema and
scope come from the catalog, not client input. Host keys come from reviewed targets.
The existing final result event contains the same object as the final response. There
is no separate live operation-data event; read final authoritative results in either
summary or detail mode.
Entry status:
| Status | Meaning |
|---|---|
| available | Complete final data received, schema/limit/safety validation passed |
| missing | Final accounting arrived but this report slot was not published |
| withheld | Sensitive publisher or known supplied secret matched output |
| invalid | Schema, provenance or size validation rejected the publication |
| not_started | Execution did not start; there is no operation data |
| indeterminate | Work may have started but no complete final report stream exists |
Unavailable entries have `data: null`; never replace them with `{}`/zero/false in a UI.
`complete` describes report availability, not execution success or payload-specific
completeness. A patch report has its own `data.complete` for update evidence. A failed
service-start operation can have an available, complete report listing failed services.
If native exit is nonzero, preserve the native/Core failure and retain any valid reports
from the final stats. If native execution exits 0 but required output is missing,
withheld, invalid or incomplete, Core returns failed/result_validation with native
exit_code 0. Native target outcomes still describe their Ansible stats. This is an
intentional additional report-contract check, not a redefinition of native stats.
Optional missing reports do not fail an otherwise successful run; invalid submitted
reports do. Never replay automatically after result-validation failure.
## Generic schema subset and safety
The schema validator is standard-library-only, shared by callback and Core. Supported
JSON Schema keywords: `type` (including nullable type arrays), `properties`, `required`,
boolean `additionalProperties`, `items`, `enum`, `maxItems`, `maxLength`, `minimum`,
`maximum`, and documentation `description`. `$ref`, arbitrary validators and executable
schema extensions are not supported. Objects are closed unless explicitly declared open.
The only shipped open subtree is already parsed/redacted Checkmk `sections`.
Limits: maximum depth 20, 200,000 nodes per value, 20,000 items per collection, 8,192
characters per string, 1 MiB encoded envelope per host/global slot and 16 MiB per run.
Catalogs can lower the per-host byte cap. Data is chunked privately into <=2,048-byte
pieces; malformed sequences, duplicate publications, torn frames and oversize data are
rejected, never silently truncated. NaN/Infinity, binary objects, cycles and unknown
closed-object properties are rejected. Schemas and output labels must be static public
controller source.
Obvious secret field names are disallowed in schemas; redacted dynamic config keys
can only hold the exact `[REDACTED]` marker. Known supplied credentials matching public
strings are withheld again at the Core boundary. These measures are defense in depth,
not a general secret scanner. Config/other strings literally containing an undisclosed
secret remain an author/operator responsibility. Never publish environments, invocation,
exception objects, arbitrary debug values, commands or private keys.
`checkmk_user_config_v1` preserves parsed sections but redacts recognized password/token/
passphrase/credential/community fields, command/argument/environment bodies, MRPE
commands, URL credentials and obvious inline secrets. `redacted_paths` explains what
was withheld. Comments/formatting are not YAML data. The operation may read only basename
`check_mk.user.yml`, <=512 KiB, from a trusted configured path. It never modifies the file.
Unknown settings remain visible unless filtered; do not advertise this as guaranteed
secret-free content. No general arbitrary-file reader is added.
## Current schemas and exact semantic boundaries
Normative field shapes: `playbooks/schemas/<schema>.yml` (also in catalog metadata).
| Schema | Data meaning |
|---|---|
| host_capabilities_v1 | Eight booleans, source detection facts; no inventory membership edits |
| filesystem_usage_v1 | Selected operational mounts/volumes; byte quantities and nullable observations; Windows uses attached storage volumes from community.windows.win_disk_facts and excludes mapped/network drives |
| event_log_export_v1 | Channel names, age and target file paths; no log contents or controller download; check mode has no exported paths |
| service_start_summary_v1 | Before stopped services, policy eligibility, attempts, exclusions, actual after observations and fixed per-attempt failures |
| patch_summary_v1 | Per-package net version-set changes on Debian/RedHat, or Windows update IDs/titles/KBs with installed flags; never arbitrary manager dictionaries |
| managed_cleanup_preview_v1 | Managed filename candidates and observed deletions; no unknown-file purge |
| checkmk_user_config_v1 | Parsed/redacted sections and file metadata, not raw YAML/debug output |
| checkmk_agent_state_v1 | Installed status/version from registry/package query, actual service states, package/config/check changes |
| checkmk_agent_config_v1 | Named section/config/check change actions; no raw before/after values or file diffs |
### Services
`initially_stopped` is the observation before action. `eligible` applies auto/delayed-start
and configured include/exclude policy. `attempted` excludes skipped and check-mode actions.
`newly_running` is the observed stopped-to-started transition, including a concurrent
external start; `started_count` counts only attempted services now running. It is not
proof that the start return alone kept a service alive. `still_stopped` includes excluded
services; `unobserved` means missing after observations. Non-running attempted services
appear in `failed_to_start` with name, fixed reason/message and nullable numeric native
code. Failed reasons include dependency_failed, permission_denied, service_disabled,
start_timeout, service_not_found, logon_failed, start_failed, not_running_after_start and
state_unavailable. Localized unrecognized errors use generic start_failed; no raw text.
An unavailable post-query fails the run instead of inventing a successful observation.
### Packages and Checkmk
Linux queries native package databases before/after the update/reboot policy, preserving
architecture and parallel installed version sets. `updates` lists updated/installed/
removed rows and separate counters. Old/new arrays are empty where the package did not
exist. Snapshots cannot identify an intermediate reinstall with unchanged final versions
or attribute concurrent external package changes; do not run another package manager
concurrently. Check mode reports no completed changes, not predicted upgrade versions.
No silently shortened list: exceeding limits becomes explicit validation failure.
Windows reports native updates marked installed, pending/failed records and numeric
failure codes. A mismatch between native installed count and detailed records marks
`data.complete: false`, rather than fabricating records. Supported scope remains Windows
Update under configured categories, not every third-party package installer. Reboot
facts retain their observed meaning; check mode never claims a performed reboot.
Checkmk versions can be null (unavailable/ambiguous); do not fall back to the filename.
Config changes are section/file-level, not field-level raw diffs. `deployed_checks` means
selected file-copy tasks completed; `changed_checks` identifies changed copy results.
Check-mode changes are predictions; mode is always explicit. Unknown files are neither
counted nor scanned: policy is `untouched_not_enumerated`, not an invented preserved count.
## Errors and consumer behavior
Report errors: operation_result_missing, operation_result_invalid,
operation_result_withheld, operation_result_limit, operation_result_incomplete.
Messages contain no payload/exception content. Preserve final Core/native status;
show report issues separately; no credential retry or task replay is implied.
Treat this candidate as an implementation awaiting native 2.19.11/managed-host acceptance.
Upstream interfaces consulted: Ansible 2.19.11 `plugins/action/set_stats.py`,
`executor/stats.py`, and callback result contracts; ansible.windows >=3.8.0,<4.0.0 win_updates and
win_service_info return definitions. Consultation is not native testing.
## Patch wave and reboot semantics (3.3.0rc8)
`patch_summary_v1` reports package/update evidence independently from reboot policy. Linux
retains the rc2 package-manager transaction model. Windows delegates each selected patch
wave to one native `ansible.windows.win_updates` `state: installed` invocation with
`reboot: false`. The collection and Windows Update Agent own update ordering/coordination
inside that wave; AIM does not schedule updates individually.
The returned update dictionary is normalized into reviewed installed/failed records. A
module-level failure with no per-update records becomes a bounded generic failed-wave record;
raw failure text, exceptions and arbitrary module dictionaries are not published.
A newly required reboot with `os_patching_reboot: false` does not invalidate successful
installation: `reboot_required`/`reboot_required_after` are true and `reboot_deferred` is
true. The run stops. On a later run, a detected pre-existing reboot with automatic reboot
still disabled publishes `evidence: preflight_reboot_state`, `complete: false`,
`blocked_reason: preexisting_reboot_required`, then fails before new patch work starts.
When automatic reboot is enabled, AIM waits for the native update wave to return before
performing its reviewed message/delay/reboot. By default
`os_patching_rescan_after_reboot: false`; any AIM-performed reboot ends the run with
`continuation_required: true` and `remaining_updates_known: false`. This includes a reboot
that was already required before patching began. A new operator-approved run owns discovery
of the next patch state.
If the operator explicitly sets post-reboot continuation true, AIM may invoke another native
Windows update wave after reboot. Continuation is bounded to 12 wave invocations per run. A
failed wave is never automatically replayed or used as authority to continue.
When a Windows wave completes without reboot, AIM performs one final read-only search. This
does not install newly applicable updates. `remaining_updates_known` is then true and
`pending` contains that final observation; `continuation_required` indicates whether another
operator-approved run has applicable work.
`patch_cycles` counts native Windows install-wave invocations entered in the run and
`rescan_after_reboot` echoes the reviewed option. Those fields are Windows-specific and are
optional in the closed schema so existing Linux result shape remains stable.
Windows `failed_updates` contains only bounded safe fields: update identity/title, normalized
unsigned HRESULT and hex form, a fixed reason and fixed message. Known reasons include
`operation_in_progress`, `install_not_allowed`, `not_applicable`,
`exclusive_install_conflict`, `self_update_in_progress`, `no_connection`, and `timeout`.
Raw Windows Update failure strings are not published. `0x80240016` maps to
`install_not_allowed`; it does not prove a reboot was pending because Windows also uses that
HRESULT while another installation is active. Only the independent preflight can report
`preexisting_reboot_required`.
`reboot_required_before` records the pre-run observation. `reboot_performed` is true when
AIM completed an approved reboot. `reboot_required_after` and the legacy
`reboot_required` field describe the final pending state represented by the report.
`reboot_delay_minutes` echoes the reviewed delay. The user-facing reboot message itself is
not copied into the structured result.
Windows, Debian and RedHat use platform-native reboot-required signals; RedHat preflight
is conservative when `needs-restarting` is not already present. Report fields never imply
that every possible vendor-specific reboot indicator was discovered.
+83
View File
@@ -0,0 +1,83 @@
# Playbooks and reporting contracts
**Current AIM 3.3.0rc8.** The catalog is authoritative for typed inputs, requirements and report declarations.
Run options omitted by a caller remain inherited from inventory/role defaults. Core
requires explicit compatible hosts and final scope review; group shortcuts expand to
deduplicated explicit hosts. `aim_debug` adds only selected diagnostics, not raw API output.
Normal terminal operations validate existing Vault passwords before launch and retain
the workflow on a typo. New-encryption/create prompts stay native. Catalog connections
use inventory/Vault credentials; automatic replay is never performed after launch.
## Catalog audit
| Key | Platforms | Report schema |
|---|---|---|
| `checkmk_install_agent` | linux, windows | `checkmk_agent_state_v1` |
| `checkmk_update_scripts_config` | linux, windows | `checkmk_agent_config_v1` |
| `checkmk_read_windows_config` | windows | `checkmk_user_config_v1` |
| `checkmk_cleanup_scripts` | linux, windows | `managed_cleanup_preview_v1` |
| `debug_test_connection` | linux, windows | None; native target outcomes/progress |
| `debug_show_disk_usage` | linux, windows | `filesystem_usage_v1` |
| `debug_detect_host_roles` | linux, windows | `host_capabilities_v1` |
| `maintenance_export_event_logs` | windows | `event_log_export_v1` |
| `maintenance_start_stopped_services` | windows | `service_start_summary_v1` |
| `maintenance_patch_os` | linux, windows | `patch_summary_v1` |
| `maintenance_reboot_hosts` | linux, windows | None; native target outcomes/progress |
| `sophos_apply_baseline` | sophosxgs | None; native target outcomes/progress |
| `sophos_apply_customer` | sophosxgs | None; native target outcomes/progress |
| `pfsense_apply_baseline` | pfsense | None; native target outcomes/progress |
Customer Sophos profiles share the catalogued customer operation; their five concrete
playbooks remain unchanged and publish no operation payload. Firewall policies, connection
test and reboot behavior were not redesigned for reporting.
## Schema and operational details
See [OPERATION_RESULTS.md](OPERATION_RESULTS.md) for field meanings and missing/partial/
check-mode behavior. Read [CHECKMK.md](CHECKMK.md) for managed filename/section ownership.
Role READMEs retain local default settings. `list_playbooks` exposes the current typed
input definitions and resolved report schemas; add-ons need not keep copies of this table.
## OS patch reporting
The catalog exposes `os_patching_reboot`, `os_patching_reboot_delay_minutes`,
`os_patching_reboot_message`, Windows categories, and the Windows-only
`os_patching_rescan_after_reboot` flag. The message/delay apply only when AIM initiates a
reboot.
Windows delegates each current patch wave to one native `ansible.windows.win_updates`
`state: installed` invocation with `reboot: false`. The module/WUA owns selection and
sequencing inside that selected category wave; AIM does not loop individual updates or use
`accept_list` to manufacture its own scheduler. Per-update records returned by the module
are normalized for reporting.
If the completed wave requires reboot, AIM either defers it or performs the reviewed
message/delay/reboot. By default (`os_patching_rescan_after_reboot: false`) that reboot ends
the run, so another wave requires a new operator-approved execution. When explicitly
enabled, AIM may start another native wave after reboot; continuation is defensively bounded
to 12 wave invocations. A failed wave is never automatically replayed.
If a wave finishes without reboot, AIM performs one final read-only discovery for reporting
only. Windows reports expose individual installed and failed updates, bounded HRESULT
classifications, `continuation_required`, and whether the reported pending list is
authoritative.
Linux snapshots use dpkg-query/RPM before and after the existing operation, retaining
architecture and parallel installed version sets. Every net difference is published;
intermediate no-net-change transactions are not observable from snapshots. No package
stdout scraping or guessed versions. Check mode never claims installation.
## Service recovery
Initial stopped/eligible/excluded states, actual attempts and post-start observations
are all separate. `failed_to_start` gives fixed reasons and numeric codes where known.
The default fail-on-error policy still fails the run after a report is published.
## Security
Catalogs, roles and inventory are trusted controller code. Report schemas are reviewed
public data contracts, not automatic secret scanners. Operator-owned Checkmk config
can contain credentials; its reader filters known secret/command fields and publishes
redacted paths. No general raw-file/data export was enabled.
+34
View File
@@ -0,0 +1,34 @@
# Current documentation index
**AIM 3.3.0rc8.** These files supersede previous release-specific handoffs/checklists.
Do not copy an old validation status into a new release. Archive history externally;
keep the reverse-chronological CHANGELOG in the distribution.
| Topic | Authoritative document |
|---|---|
| Release changes | [RELEASE_NOTES.md](RELEASE_NOTES.md) |
| Current evidence and limits | [VALIDATION.md](VALIDATION.md) |
| Operator acceptance procedure | [SANITY.md](SANITY.md) |
| Fresh installation | [INSTALLATION.md](INSTALLATION.md) |
| Update and recovery | [deploy/README.md](../../deploy/README.md) |
| Core development | [AGENTS.md](AGENTS.md) |
| Add-on implementation responsibilities | [ADDON_AGENTS.md](ADDON_AGENTS.md) |
| API request/response lifecycle | [ADDON_API.md](ADDON_API.md) |
| Purposeful data and schema authoring | [OPERATION_RESULTS.md](OPERATION_RESULTS.md) |
| Safe execution progress | [DETAILED_PROGRESS.md](DETAILED_PROGRESS.md) |
| Native target outcomes | [TARGET_OUTCOMES.md](TARGET_OUTCOMES.md) |
| Inventory tree | [INVENTORY_HIERARCHY.md](INVENTORY_HIERARCHY.md) |
| Hardened execution environment | [EXECUTOR_STAGING.md](EXECUTOR_STAGING.md) |
| Implemented/deferred boundary | [ADDON_SUPPORT.md](ADDON_SUPPORT.md) |
| Current handoff checklist | [RELEASE_HANDOFF.md](RELEASE_HANDOFF.md) |
| Playbook options and reports | [PLAYBOOKS.md](PLAYBOOKS.md) |
| Checkmk-specific operational settings | [CHECKMK.md](CHECKMK.md) |
`addon-support-v1.json` is a snapshot of default capabilities, not a release manifest
or runtime acceptance certificate. `aimctl capabilities` reports the installed state.
The nine files in `playbooks/schemas/` are normative data shapes; schema definitions
are also resolved into catalog metadata so add-ons need not access the filesystem.
- [CHANGELOG.md](CHANGELOG.md) — retained release history.
- `roles/*/README.md` — role-specific reference documentation stays beside each role.
- `deploy/README.md` — deployer/update/rollback mechanics stay beside the deployer.
+91
View File
@@ -0,0 +1,91 @@
# AIM 3.3.0rc8 Core and add-on handoff
- AIM-owned documentation is centralized under `scripts/docs/`; the installation-root `README.md` is operator-owned and preserved.
## 3.3.0rc8 final Windows Checkmk script placement
- `citrix_sessions_customized.ps1`, `veeam_o365_status.ps1`, and `veeam_backup_status.ps1` are AIM-managed custom plugins under `$CUSTOM_PLUGINS_PATH$`.
- `veeam_backup_license_status.ps1` is a VBR-detected local check under `$CUSTOM_LOCAL_PATH$`.
- AIM emits exact custom-plugin rules before the generic plugin rules and no longer needs to place these custom scripts in Checkmk's built-in plugin tree.
- Selected custom-plugin deployment removes only known legacy AIM-managed copies from the historical local directory and, where applicable, the historical built-in plugin directory. Unknown/custom files remain untouched.
- Service/wire/event API 1.0 and the existing Checkmk structured result schemas are unchanged.
## Release status
Candidate implementation is complete for the nine agreed reporting operations. Local
validation and limitations are maintained only in [VALIDATION.md](VALIDATION.md).
No new native controller/Windows/Linux acceptance is claimed. The remaining operator
gate is [SANITY.md](SANITY.md). Do not relabel older successful runs as candidate tests.
## What independent teams consume
The stable boundary remains `aim.services.v1` / `aimctl`, not private callbacks or UI
modules. Inspect `operation_results`, then catalog `result`/prepared `result_contract`.
Consume final `operation_result` in either progress mode. No add-on source was inspected,
modified or required. Existing adapters continue to receive previous fields; an updated
adapter renders the additional data without reconstructing it from task events.
Read the authoritative [add-on guide](ADDON_AGENTS.md), [API](ADDON_API.md),
[result protocol](OPERATION_RESULTS.md) and [support matrix](ADDON_SUPPORT.md). Exact
payload shapes live in `playbooks/schemas/` and are included in public catalog metadata.
Keep overall Core status, native per-target outcomes and report availability separate.
A native exit 0 plus missing required output yields Core result_validation failure.
A native failed job can still provide useful report data. Do not replay either case
without an explicit new operator decision, approval and fresh credentials.
## Implementation map for future Core work
- `integrations/output_policy.py`: standalone JSON schema subset and data validator.
- `playbooks/catalog.py`: resolve/validate catalog-owned schemas.
- `integrations/callbacks/aim_safe_events.py`: verify set_stats provenance/no_log,
reconcile final custom stats and send bounded private chunks.
- `runtime/outputs.py`: reassemble/revalidate, apply known-secret suppression and
construct final report availability. Independent of detail/summary mode.
- `services/v1`: optional metadata/final result fields and unchanged lifecycle controls.
- `playbooks/filter_plugins/aim_reports.py`: runbook-owned normalizers, no Core dispatch.
- `playbooks/schemas`: nine static data shapes. Add new schemas without Core branches.
## Deployment
Use the full replacement ZIP and checksum with the normal operational deployer. Prefer
automatic existing-interpreter discovery; do not require a remembered AIM_PYTHON variable.
Review retirement/removal entries as well as writes. Quiesce before apply and preserve
recovery data. The narrow executor staging exceptions remain necessary. Core does not
modify unit files or prove another service account can use its paths/keys/collections.
## Open qualification and limits
Native set_stats/callback behavior on 2.19.11, real Windows/Linux reports and package
manager tests remain controller gates. Global transport is implemented but no bundled
operation uses it; native global/serial-run_once behavior needs separate qualification.
Config filtering is an explicit safety subset, not a universal secret scanner. Package
snapshots describe net observed changes, not an exhaustive transaction journal. No raw
result API, Custom credential overrides, key export or automatic permission migration.
## 3.3.0rc8 native-module baseline
The canonical controller remains ansible-core 2.19.11, but Windows playbooks now require ansible.windows >=3.8.0,<4.0.0 so Core can use the collection-owned `win_reboot_info` detector. Existing 3.2.x installations must update that collection explicitly before rc5 acceptance. `aimctl capabilities` advertises the floor and readiness rejects older versions.
The audit also moved Checkmk file reads to `win_stat`/`slurp`, Linux package snapshots/version queries to `package_facts`, and Linux Checkmk final service reporting to `service_facts`. Reviewed custom commands remain only where current native modules do not preserve AIM's required semantics. See RELEASE_NOTES.md.
## 3.3.0rc8 patch-runbook delta
The public Core API remains 1.0. `os_patching_rescan_after_reboot` and the additive Windows
fields in `patch_summary_v1` remain catalog/result-driven. Do not hard-code UI defaults.
rc4 removes rc3's AIM-owned per-update queue. Each Windows patch wave is one native
`ansible.windows.win_updates` install invocation with the selected categories and
`reboot: false`. This intentionally delegates update ordering/coordination inside the wave
to the collection and Windows Update Agent while preserving AIM ownership of the reviewed
reboot message, delay, and whether another post-reboot wave is authorized.
The default continuation remains false. After any AIM-performed reboot, including a
pre-existing-reboot preflight, the run stops unless the operator explicitly enabled
post-reboot continuation. A failed wave is never automatically replayed. Explicit
continuation remains defensively capped at 12 waves.
Per-update records returned by `win_updates` are still normalized into bounded installed and
failed records. `0x80240016` remains `install_not_allowed`, not proof of a reboot; the
independent preflight remains the source for `preexisting_reboot_required`.
+72
View File
@@ -0,0 +1,72 @@
# AIM 3.3.0rc8 release notes
3.3.0rc8 is a focused Windows Checkmk ACL hardening release on top of the final rc7 script-placement baseline. AIM-managed persistent Checkmk files now remain readable/manageable by the standard local administrative principals even when created by a service account.
## Windows Checkmk managed-file ACLs
After AIM creates or updates a persistent Windows Checkmk script or `check_mk.user.yml`, the reusable `checkmk_windows_acl` role enables parent ACL inheritance and guarantees these locale-independent well-known SID entries:
- SYSTEM (`S-1-5-18`): FullControl
- local Administrators (`S-1-5-32-544`): FullControl
- ALL APPLICATION PACKAGES (`S-1-15-2-1`): ReadAndExecute
- ALL RESTRICTED APPLICATION PACKAGES (`S-1-15-2-2`): ReadAndExecute
AIM does not add a customer-specific administrator/user ACE such as the example `bitformer` account. ACL normalization is per exact AIM-managed file; it does not recurse through Checkmk directories or modify unknown/operator files. Existing intentional inherited/explicit ACEs are not blindly purged.
## Final Windows Checkmk script placement
`citrix_sessions_customized.ps1`, `veeam_o365_status.ps1`, and `veeam_backup_status.ps1` are now consistently deployed as Checkmk custom plugins under `C:\ProgramData\checkmk\agent\plugins` (`$CUSTOM_PLUGINS_PATH$`). Their managed execution rules use the custom-plugin path; AIM no longer needs to overwrite Checkmk's built-in plugin tree. `veeam_backup_license_status.ps1` is added as a VBR-detected local check under `$CUSTOM_LOCAL_PATH$`. Selected plugin deployment removes only known historical AIM copies from the old local/built-in locations; unknown files remain untouched.
## Windows disk facts
- `debug_show_disk_usage` now uses `community.windows.win_disk_facts` instead of a custom `Get-PSDrive` PowerShell collector.
- Windows results represent attached local volumes. Mapped/network drives are intentionally outside this host-capacity report.
- `filesystem_usage_v1` is unchanged, so existing add-on renderers do not need a schema migration.
Baseline: complete AIM 3.3.0rc8 replacement bundle. Service/wire/event API remains 1.0 and canonical Ansible Core remains 2.19.11.
## Native-module audit
This candidate reviews the complete bundled playbook/role tree with a native-module-first rule: if the supported Ansible runtime already exposes the required semantics, AIM delegates to that module instead of maintaining its own shell/PowerShell implementation.
Changes made:
- Windows pending-reboot detection now uses `ansible.windows.win_reboot_info` and publishes its bounded reboot sources instead of AIM-maintained registry heuristics.
- `checkmk_read_windows_config` now uses `ansible.windows.win_stat` plus `ansible.windows.slurp`; no PowerShell is used to stat/read the file.
- Debian and RedHat package before/after snapshots use `ansible.builtin.package_facts` instead of direct `dpkg-query`/`rpm -qa` commands.
- Linux Checkmk installed-package reporting uses `package_facts`; Linux Checkmk runtime state reporting uses `service_facts` instead of `dpkg-query`/`rpm` and `systemctl show` for the final report.
- Core/service and terminal dependency preflights now reject `ansible.windows` older than 3.8.0 before a playbook can rely on `win_reboot_info`.
## Collection baseline change
`ansible.windows.win_reboot_info` was introduced in ansible.windows 3.8.0. The release requirements therefore declare:
```yaml
- name: ansible.windows
version: ">=3.8.0,<4.0.0"
```
Existing controllers that still have ansible.windows 3.2.x must deliberately update the collection before accepting rc8. AIM still does not install or upgrade collections automatically.
## Reboot-state behavior
The Windows preflight now trusts the collection's reboot detector, which covers Windows Update, Component Based Servicing, pending file rename, pending computer rename, domain join and Server Manager sources. The public patch report can include `reboot_reasons_before` entries containing bounded `source` and `description` fields.
Overall patch-wave/reboot policy is unchanged from rc4: one native `win_updates` wave, AIM-controlled reboot message/delay, and no post-reboot patch wave unless `os_patching_rescan_after_reboot` was explicitly enabled.
## Reviewed custom operations intentionally retained
The audit did not replace custom code where a native module would lose required behavior:
- Windows disk reporting now uses `community.windows.win_disk_facts` and reports attached local storage volumes. Mapped/network drives are intentionally excluded from Windows host-capacity reporting.
- Windows event-log export keeps a bounded `win_powershell` wrapper around the native Windows event export utility because the supported collection does not provide EVTX export with the required time filter.
- Checkmk's Windows `plugins:` section editor remains custom because AIM must preserve every unmanaged section/comment while replacing only its marked section.
- Windows Checkmk installed-version discovery retains a bounded read-only uninstall-registry query because no supported module exposes arbitrary installed MSI/application inventory with the required product matching semantics.
- RedHat reboot-required checks retain `needs-restarting -r`; no supported ansible-core module exposes that host state.
- Linux Checkmk unit discovery retains `systemctl show LoadState` because AIM supports socket activation and `service_facts` is service-oriented and is not a reliable replacement for arbitrary socket-unit existence discovery.
- pfSense's two `raw sysctl` operations remain intentionally bootstrap-safe and do not add a Python/runtime dependency merely to replace two appliance-native calls.
This is therefore a native-first policy, not a prohibition on all commands.
All AIM-owned Markdown/JSON documentation is centralized under `scripts/docs/`. The release no longer owns the installation-root `README.md`; an existing root README is preserved during update.
+105
View File
@@ -0,0 +1,105 @@
# AIM 3.3.0rc8 controller acceptance
Run on disposable/approved targets under the actual execution UID, groups, HOME and
service sandbox, using native ansible-core 2.19.11 and declared collections. Do not use
an ordinary root shell as evidence for another worker. Do not send raw Vaults/config
secrets in test feedback. Keep the staging exceptions from the accepted deployment.
## Installation and unchanged controls
Verify checksum, preview the deployer, confirm only named obsolete Core docs are removed,
quiesce jobs/writers and apply. Verify `aim --version`, `aimctl --version`, capabilities,
independent terminal startup, Vault wrong-then-correct retry and cancellation. Confirm
operator aim.yml/inventory/keys/add-on state are unchanged. `aimctl staging-check` must
pass inside the executor, including differing process and passwd homes when relevant.
## Discover the new contract without credentials
```bash
aimctl capabilities
printf '%s\n' '{"api_version":"1.0","operation":"list_playbooks","customer":"CUSTOMER"}' | aimctl request
printf '%s\n' '{"api_version":"1.0","operation":"prepare","request":{"customer":"CUSTOMER","playbook":"debug_detect_host_roles","hosts":["HOST"]}}' | aimctl request
```
Replace identifiers. Request JSON is one line. Expect operation_results capability and
prepared result_contract with host_capabilities_v1; unchanged explicit targets/revision.
Nonreporting `debug_test_connection` must advertise result:null. Unknown options/hosts
must still fail before launch. Schema changes must stale an earlier preparation.
## Fresh approved execution matrix
Supply credentials through the established one-run provider/private FD. Use summary
and detail on separate newly approved runs; compare reports rather than event ordering.
Do not cache the password or reuse a provider/revision after source changes.
| Operation | Evidence required |
|---|---|
| Detect roles | Eight booleans agree with authorized native report; no inventory mutation |
| Disk usage | Compare Windows/Linux observed byte counts; handle unavailable/empty mounts |
| Export event logs | Paths exist after apply; existing logs not cleared; check creates no export |
| Start services | Disposable services: one recoverable, one failing dependency/disabled/race case, one excluded; compare before/after states and fixed reasons |
| Patch OS | Disposable snapshots: verify Linux net package/version changes; Windows uses one native win_updates wave per approved patch stage; IDs/KBs/HRESULTs; user-visible reboot message/delay; default stops after reboot with continuation_required; explicit post-reboot continuation starts another wave |
| Checkmk cleanup | Preview removes nothing; approved deletion lists only managed files actually changed |
| Read Checkmk user config | Plugins/local/unknown sections preserved in report; dummy passphrase and MRPE command data redacted; reject other basenames; no timestamp/content write |
| Checkmk install | Installed version agrees with registry/package DB; services observed; idempotent rerun shows no package/config changes |
| Checkmk config update | One approved rule/file change reported; second run no change; unrelated local/MRPE/unknown files untouched |
New reporting tasks change task counts. Compare effects, failure states and payloads,
not historical exact ok/skipped totals. For patching, protect against concurrent package
management and record starting snapshot/selected categories. A partial update still
requires explicit recovery decisions; never repeat automatically.
### Windows patch-wave acceptance
Use a disposable Windows target with several applicable updates. Confirm the normal apply
path contains one `win_updates state=installed` task for the selected categories rather than
one task per update. Windows/collection-native sequencing inside that wave is expected; AIM
should not emit `accept_list` selectors for individual discovered updates.
Record the structured result and compare it to Windows Update history. Installed and failed
updates should retain bounded IDs/titles/KBs/HRESULT classifications. A mixed native result
must not lose successful updates merely because another update failed. Raw failure messages
or exception text must not enter the operation result.
With `os_patching_rescan_after_reboot: false` (catalog default), allow the current wave to
finish and require reboot. The user should receive the reviewed message/delay. After
reconnect the run must end without another install/search wave. Expect
`continuation_required: true` and `remaining_updates_known: false`; a new approved run owns
the next patch state.
Also test a host that begins with an already-pending reboot. With automatic reboot enabled
and post-reboot continuation false, AIM should reboot and stop before patch installation.
With continuation true, it may begin the first native update wave after reboot. With
automatic reboot disabled, it must fail preflight before new installs.
Repeat with `os_patching_rescan_after_reboot: true`. After a patch-triggered reboot AIM may
start another native update wave. Do not infer this opt-in from the reboot flag. Explicit
continuation remains bounded to 12 waves. A failed wave must stop and must never be replayed
automatically.
For a failed update, confirm `failed_updates` contains only bounded title/ID, unsigned/hex
HRESULT, reason and fixed message. If `0x80240016` can be reproduced, expect
`install_not_allowed`; do not expect `preexisting_reboot_required` unless the separate
preflight actually observed a pending reboot.
## Fail-closed fixtures (disposable catalog/playbook only)
After each fixture edit, prepare/review anew. Test missing required publisher; wrong
schema/type/extra fields; publisher under no_log; duplicate publication; >limit data;
and a supplied secret canary in an output string. Expect no rejected payload in public
JSON/logs. Native exit 0 with invalid/missing required data must fail at result_validation,
retain exit_code 0/native target facts and never replay. No declaration must export
nothing even if unrelated debug/custom stats contain a secret.
Test two targets with one unreachable and one completed report. Overall native failure
remains; one available report does not imply all-target success. Test Ctrl+C mid-run:
report must not claim earlier observations are completed output. Ensure no residual
credential/helper process and a subsequent fresh operation works. Test chunked output
beyond a single private frame; a torn stream must fail, never silently truncate.
## Sign-off
Record version, execution UID/group context, runtime/collection versions, platform and
which cases were tested, separately for native terminal and service. Do not mark global
transport, another OS, Server 2012 R2 or another service identity accepted from one Windows
11 test. Report final status/exit/target facts/report availability with credentials removed.
+108
View File
@@ -0,0 +1,108 @@
# AIM per-target final execution outcomes
**AIM 3.3.0rc8 / service API 1.0. Capability:** `target_outcome_summary_v1`.
Core's overall execution result remains authoritative and unchanged. An Ansible run
with one unreachable target still returns, for example, `status: failed`, execution
stage and the native nonzero exit code. The additive target summary answers a
different question: what final Ansible outcome did each **requested target** have?
## Result fields
Every `RunResult` now contains:
```json
{
"target_summary": {
"schema": "target_outcome_summary_v1",
"requested": 25,
"successful": 24,
"failed": 0,
"unreachable": 1,
"not_started": 0,
"indeterminate": 0,
"complete": true,
"accounted": 25
},
"targets": [
{
"host": "host01.example",
"outcome": "successful",
"counts": {
"ok": 38,
"changed": 0,
"failures": 0,
"unreachable": 0,
"skipped": 22,
"rescued": 0,
"ignored": 0
}
}
]
}
```
The `targets` list is in the reviewed request order. `accounted` always equals
`requested`. Capability discovery advertises the schema, state vocabulary and that
the data is available in both `summary` and `detail` progress modes.
## Outcome semantics
Core derives these states from Ansible's **final per-host stats callback**, not from
aggregate task counters or a client's interpretation of progress events:
- `unreachable`: final host stats contain one or more unreachable results.
- `failed`: no unreachable count, but final host stats contain unresolved failures.
- `successful`: final host stats were emitted for the requested host and contain
neither unresolved failures nor unreachable results. Changed/skipped/rescued/
ignored task counts do not by themselves make the host unsuccessful.
- `not_started`: a complete final stats set was received, but the requested host
had no final host stats; or execution ended before remote work could start.
- `indeterminate`: remote execution may have started but Core did not receive a
complete final stats set (for example cancellation, process loss or invalid
event stream). Core deliberately does not infer success from earlier task events.
If both failure and unreachable counters exist for one host, `unreachable` takes
precedence because the target did not remain reachable through the execution.
Per-target `counts` are task counters for diagnostics; the `outcome` field is the
Core-owned final classification.
This describes Ansible execution truth, not arbitrary application-level semantics.
A playbook that intentionally ignores a module error or reports a domain-specific
problem without failing remains subject to the playbook's own Ansible semantics.
## Overall status remains separate
A mixed run can therefore be represented as:
```text
Core RunResult: failed / execution / exit 4
Target facts: 24 successful, 1 unreachable
```
An interface may choose a presentation label such as `Partially succeeded`, but
that label is not a Core status and must not replace or hide the authoritative
Core result.
Pre-execution `ServiceError` responses remain errors rather than fake partial
success. When `execute` returns a `RunResult` before launch, targets are accounted
as `not_started` and `remote_work_may_have_started` remains false.
## Transport and safety
The native callback sends only target indexes and numeric recap counts through the
private event bridge. Public host names come from the already reviewed request,
not arbitrary callback payloads. No raw module output, variable data, credentials,
exception text or task result dictionaries are added.
Consumers should use this result instead of reconstructing final target outcomes
from `play_task_host_v1`. Detailed events remain useful for live presentation; the
final target summary is the authoritative end-state accounting.
## Report validation is a separate layer
A native exit 0 and successful target stats do not prove a required operation report
was published. 3.3.0rc8 can fail the Core result at result_validation while preserving
those native target facts and exit_code 0. See OPERATION_RESULTS.md. Do not convert a
missing report into fake target success, nor rewrite native counts to express a
report-contract error.
+64
View File
@@ -0,0 +1,64 @@
# AIM 3.3.0rc8 validation record
This record distinguishes source/local validation from native controller acceptance.
## Candidate scope
3.3.0rc8 is a focused Windows Checkmk ACL hardening release on top of the final rc7 source state. It retains the native-module, disk-facts, patch-wave, Checkmk placement, and public API behavior while normalizing access on AIM-managed persistent Windows Checkmk files.
## Local checks
Passed for this source tree:
- all bundled Python files parsed successfully;
- all YAML/YML files parsed successfully;
- report filters compile and the package-facts normalizer accepts Debian/RPM-shaped fact dictionaries;
- `patch_summary_v1` accepts the additive bounded `reboot_reasons_before` field;
- Windows patch source contains `ansible.windows.win_reboot_info` and no custom reboot-registry PowerShell probe;
- Checkmk config reading contains `win_stat` + `slurp` and no `win_shell` reader;
- Linux patch package snapshots contain `package_facts` and no `dpkg-query`/`rpm -qa` snapshot command;
- Linux Checkmk final reporting contains `package_facts`/`service_facts` instead of direct package/systemctl queries;
- Core/service and terminal preflight reject an ansible.windows version below 3.8.0.
- Windows disk usage uses `community.windows.win_disk_facts`; the PowerShell `Get-PSDrive` collector is absent. A local normalizer fixture verified drive-letter and no-drive-letter attached volumes against the unchanged `filesystem_usage_v1` shape.
- the new `checkmk_windows_acl` role uses `win_acl_inheritance` plus `win_acl` with well-known SIDs and is called only for exact AIM-managed script/config paths;
- no customer-specific principal name such as `bitformer` is present in the ACL role;
- no recursive ACL task targets the whole Checkmk local/plugins/config directories.
The remaining command/PowerShell/raw call sites were individually reviewed and retained only where current supported modules do not preserve AIM's required semantics; see RELEASE_NOTES.md.
A wheel build with `pip wheel --no-deps --no-build-isolation ./scripts` succeeded. After removing generated build residue, a disposable fresh full-source installation succeeded with both installed launchers reporting AIM 3.3.0rc8. A disposable rc7→rc8 replacement update also succeeded while preserving an operator root README, unknown operator documentation, and operator `scripts/aim.yml`.
## Exact-runtime limitation
This build environment cannot install the canonical external Ansible runtime/collections from package repositories, so no native `ansible-playbook --syntax-check`, `win_reboot_info`, WinRM, package-manager, or managed-host execution is claimed here.
The release therefore requires controller qualification with:
- ansible-core 2.19.11;
- ansible.windows >=3.8.0,<4.0.0;
- the remaining collections from requirements.yml;
- an approved Windows target and Linux package-manager targets as applicable.
## Required controller acceptance
Before stable promotion verify at minimum:
1. `ansible-galaxy collection list ansible.windows` reports 3.8.x and AIM readiness succeeds;
2. a Windows host with no pending reboot reports `reboot_required_before=false`;
3. a Windows host with a real pending reboot reports the native reason(s), and false-positive behavior is improved relative to the old registry probe;
4. the current Windows update wave/reboot/no-rescan policy from rc4 remains unchanged;
5. `checkmk_read_windows_config` returns the same bounded/redacted report using stat/slurp;
6. Debian/RedHat patch reports still list actual net package changes using package_facts snapshots;
7. Checkmk Linux install/update reporting still returns installed version and service state correctly;
8. a Windows Checkmk deployment created by the service account leaves each AIM-managed script and `check_mk.user.yml` readable/manageable by local Administrators and SYSTEM and readable/executable by both application-package principals.
Previous rc4 Windows patch success is historical evidence only and is not relabeled as an rc8 test.
## Final Windows Checkmk placement checks
Local release validation confirms the Windows script catalog marks `citrix_sessions_customized.ps1`, `veeam_o365_status.ps1`, and `veeam_backup_status.ps1` for `$CUSTOM_PLUGINS_PATH$`, while `veeam_backup_license_status.ps1` is selected automatically as a local check when VBR is detected. The rendered plugin template targets the exact custom-plugin paths, and migration/cleanup touches only the documented AIM-managed legacy filenames. Documentation-source validation also confirms no release Markdown/JSON remains outside `scripts/docs/`, and update planning leaves the installation-root `README.md` untouched. Native Windows/Checkmk execution still requires controller acceptance.
## Documentation placement
Validated that only root/scripts-root AIM-owned documentation is centralized under `scripts/docs/`; `deploy/README.md` and role READMEs remain component-local. The root `README.md` is not part of the release payload and remains operator-owned during update planning.
+135
View File
@@ -0,0 +1,135 @@
{
"api_version": "1.0",
"event_version": "1.0",
"core_version": "3.3.0rc8",
"canonical_ansible_core": "2.19.11",
"contract_stability": "stable_1.x",
"credential_requirements_policy": "conservative_customer_vault",
"operations": [
"capabilities",
"list_customers",
"list_hosts",
"inventory_hierarchy",
"list_playbooks",
"prepare",
"readiness",
"execute",
"staging_check"
],
"execution": {
"implemented": true,
"enabled": false,
"profile": "same_uid_native_inventory",
"qualification": "controller_acceptance_required"
},
"execution_progress": {
"modes": [
"summary",
"detail"
],
"default": "summary",
"request_field": "progress_mode",
"detail_schema": "play_task_host_v1",
"detail_event_kinds": [
"play_started",
"play_skipped",
"play_stopped",
"task_started",
"host_result",
"task_retry",
"task_async_poll",
"host_recap"
],
"static_source_labels_only": true,
"raw_output": false,
"error_policy": "fixed_diagnostic_hints",
"qualification": "controller_acceptance_required"
},
"inventory_hierarchy": {
"schema": "inventory_hierarchy_v1",
"read_only": true,
"nested_groups": true,
"direct_hosts_only_per_node": true,
"execution_targets_remain_explicit_hosts": true
},
"target_outcomes": {
"schema": "target_outcome_summary_v1",
"states": [
"successful",
"failed",
"unreachable",
"not_started",
"indeterminate"
],
"source": "native_final_host_stats",
"available_in_progress_modes": [
"summary",
"detail"
],
"overall_status_unchanged": true
},
"operation_results": {
"protocol": "aim_output_v1",
"result_protocol": "aim_operation_result_v1",
"catalog_field": "result",
"scopes": [
"per_host",
"global"
],
"available_in_progress_modes": [
"summary",
"detail"
],
"raw_output": false,
"max_bytes_per_host": 1048576,
"max_bytes_per_run": 16777216,
"qualification": "controller_acceptance_required"
},
"controller_staging": {
"profile": "native_defaults_preflight_v2",
"operation": "staging_check",
"default_directory": "~/.ansible/tmp",
"passwd_home_local_connection_directory": "~account/.ansible/tmp",
"required_before_credentials": true,
"rechecked_before_launch": true,
"probe": "create_write_read_remove",
"service_units_modified": false,
"remote_paths_overridden": false,
"coverage": "native_local_tmp_and_passwd_home_posix_local_default"
},
"credential_fields": [
"vault_password",
"connection_password",
"become_password",
"ssh_key_passphrase"
],
"credential_semantics": "native_defaults_not_inventory_overrides",
"unsupported": [
"cross_uid_execution",
"custom_password_override",
"private_key_export",
"inventory_mutation",
"vault_mutation",
"bootstrap",
"raw_task_output",
"automatic_permission_migration",
"automatic_retry_after_launch",
"untrusted_tenant_sandbox"
],
"limits": {
"request_bytes": 131072,
"credential_bytes": 65536,
"hosts": 1000,
"timeout_seconds": 86400,
"credential_ttl_seconds": 300,
"detail_events": 200000,
"progress_label_characters": 200,
"progress_host_characters": 255
},
"document_type": "default_capability_snapshot",
"qualification_record": "VALIDATION.md",
"contract_index": "README.md",
"collection_baselines": {
"ansible.windows": ">=3.8.0,<4.0.0"
}
}