Skip to content
THE LIGHTHOUSE // RECOVERY
Observatory Self-Healing

Detect. Recover. Verify.

Bring monitoring, investigation, and recovery into a connected operational workflow. Observatory helps teams respond through approved actions, service-aware policies, and clear evidence of recovery.

  1. DETECTA dependency fault is picked up by monitoring.
  2. DIAGNOSEThe affected path is traced and investigated.
  3. AUTHORIZERecovery policy decides what may run, and who approves.
  4. RECOVERThe approved workflow runs.
  5. VERIFYChecks confirm the service works again.
INVESTIGATIONEVIDENCE LINKED
  • Likely cause · dependency degraded
  • Impact · Orders API → Customer portal
  • Evidence · correlated signals and changes
RECOVERY POLICYAWAITING APPROVALAPPROVED
  • Service · production
  • Mode · approval required
  • Approver · authorized operator
RECOVERY ACTIONRUNNINGACTION COMPLETE
  • Workflow · approved restart of the dependency
VERIFICATIONCHECKINGSERVICE STABLE
  • Resource health
  • Application check
  • Synthetic check
  1. A dependency develops a visible fault.
  2. The lighthouse beam reveals the affected service path.
  3. An investigation card links the evidence.
  4. A recovery-policy checkpoint appears: approval is required.
  5. The approved action progresses.
  6. Verification checks complete.
  7. The service returns to a stable state.

ILLUSTRATIVE RECOVERY WORKFLOW · CONCEPTUAL, NOT A PRODUCT RECORDING

A complete recovery journey

See every step toward service restoration.

  1. DetectMonitoring picks up operational symptoms across services, applications and infrastructure.
  2. InvestigateSignals are correlated and likely causes investigated along the service's dependencies.
  3. SelectAn applicable recovery workflow is selected for the affected resource and service.
  4. AuthorizePermissions and approval requirements are applied before anything changes.
  5. ExecuteThe authorized response runs. If a step fails, compensating steps can undo the work already done.
  6. VerifyChecks confirm functionality, and the service is observed for stability.
  7. RetainThe run, its approvals and its evidence are kept, and recovery knowledge is reviewed before reuse.
ACTION COMPLETEThe workflow finished its steps. That alone does not mean the service works.
≠
SERVICE RECOVEREDVerification checks passed and the service is stable. Only then is recovery recorded as successful.
Choose your level of automation

Automation that follows your operating policy.

OBSERVENO CHANGES MADE

Collect evidence and identify recovery opportunities.

RECOMMENDA PERSON DECIDES

Prepare a proposed response for review.

APPROVERUNS AFTER APPROVAL

Execute after authorized approval.

AUTOMATEWITHIN DEFINED LIMITS

Run approved recovery workflows within defined limits.

Set different policies for development, testing, and production services. Keep operators informed about what is permitted, what is running, and what needs attention.

Some limits are fixed. Destructive actions always wait for a person, even under automation. When a person requests a recovery run, a different person must approve it.

Recovery built around services

Understand dependencies before making changes.

Infrastructure actions can affect multiple applications and teams. Connect recovery decisions to service dependencies, available capacity, and the impact of concurrent changes.

Observatory business service view — technical services rolled up under each business service, with alerts, incidents, impact and SLA state

PRODUCT SCREENSHOT — BUSINESS SERVICES, THE SERVICE CONTEXT THAT RECOVERY DECISIONS DRAW ON

Human and AI collaboration

Use proven workflows. Bring in AI when investigation needs more help.

Known incidents can follow approved recovery workflows. For unfamiliar problems, AI-assisted investigation can help analyze evidence and propose next steps within configured permissions.

New recovery knowledge is reviewed before it becomes an approved automatic response.

How that works today: a workflow drafted with AI starts disabled, and a person has to promote it before it can run. AI proposals go through the same authorization and execution controls as any other recovery plan.

Recovery visibility

Every recovery, in plain view.

Operators can see each recovery from start to finish in the Observatory console. We show these screens live rather than as mock-ups.

Active recovery operationsRuns in progress, and which step each one is on.
Investigation timelineThe evidence, correlation and reasoning behind a proposed response.
Approval requestsWhat is waiting for a decision, and who can approve it.
Workflow progressStep-by-step state, including compensating steps.
Verification resultsWhich checks passed or failed after the action.
Recovery historyPast runs, approvals and decisions, kept for review.

CONCEPTUAL VIEW · NOT A PRODUCT SCREENSHOT

Infrastructure and AI workflows

One recovery model, two kinds of estate.

Infrastructure and applications

  • Service failures
  • Dependency issues
  • Configuration problems
  • Approved operational recovery

AI workflows IN DEVELOPMENT

  • Model endpoint failures
  • Tool errors
  • Stalled flows
  • Policy-approved recovery paths through AFPM
Any model change must follow configured policy and remain visible to the operator.

Available now: AFPM traces AI flows and shows where they fail. Coming next: recovery for AI flows, built on the same policy, approval and verification controls as infrastructure recovery. About AFPM →

Fit into your environment

Recovery that works where your estate runs.

Recovery actions run through Spectient, Tracston's agent that runs next to your systems, or through supported agentless integrations such as SSH, WinRM, Kubernetes, Docker, Proxmox and AWX.

Available recovery actions depend on the connected platform, installed capabilities, and granted permissions.

Observatory looks after itself too. A built-in watchdog finds and repairs faults in Observatory's own collectors and queues. It is not allowed to delete data.

Questions

Self-Healing FAQ

Does every recovery require AI?

No. Known recovery workflows run deterministically, without an LLM. AI-assisted investigation can be used when it is enabled and appropriate. It works through defined tools within configured permissions, and in production it is limited to recommending a response.

Can production changes require approval?

Yes. Recovery policy can require approval and restrict which actions are permitted for production services. Destructive actions always wait for a person, whatever the automation level. When a person requests a recovery run, a different person must approve it.

How is recovery verified?

Verification can combine resource health, service and application checks, metrics and service-level objectives, and synthetic checks against user-facing endpoints, depending on the configured service. A workflow step that completes is not treated as recovery until those checks pass.

Can Observatory work without Imperium?

Yes. Observatory includes its own recovery controls: policies, approvals, verification and audit history. Imperium is a separate Tracston product and is not required for recovery. The two connect when both are deployed.

Does recovery permanently fix every problem?

No. Some actions restore availability while a permanent correction still needs investigation. Recovery workflows can open a ticket so that follow-up work is recorded, and lessons from repeated recoveries feed back into recovery policy.

Capabilities are enabled according to deployment and licensing. Recovery depends on your configuration and environment; no recovery outcome is guaranteed.