Skip to content

DCM User Flows

This document summarizes the primary user flows in the DCM system, covering policy management, service type and catalog item management, service provider and agent lifecycle, and end-to-end CatalogItemInstance creation and deletion.

Table of Contents


1. System Overview

The DCM system is composed of the following core components:

ComponentResponsibility
Catalog ManagerEntry point for user requests; manages CatalogItems and CatalogItemInstances
Catalog DBStores CatalogItems, CatalogItemInstances, and ServiceType definitions
Placement ManagerOrchestrates instance creation; coordinates policy evaluation and agent selection
Policy Manager (Policy Engine)Validates, mutates, and selects Agents via REGO policies and OPA
SP Resource ManagerIntermediary between Placement Manager and Agents; publishes CloudEvents to agent topics; consumes responses
Agent RegistryStores Agent registration data (name, environment, service_types, topic_name, cost, health_status)
Service ProvidersExecute infrastructure provisioning (KubeVirt SP, K8s Container SP, K8s Storage SP, ACM Cluster SP)
Environment AgentRuns in target environment; routes creation/deletion requests to SPs; monitors SP health; reports to DCM via heartbeats
Messaging SystemHandles CloudEvents for asynchronous request delivery and status reporting (NATS)
    graph TB
    User([User])
    Admin[Admin]
    CM[Catalog Manager]
    CDB[(Catalog DB)]
    PM[Placement Manager]
    POL[Policy Manager / OPA]
    SPRM[SP Resource Manager]
    AR[(Agent Registry)]
    DB[(Placement DB)]
    MS[Messaging System / NATS]

    AG1[Agent - Environment 1]
    AG2[Agent - Environment 2]
    SP1[KubeVirt SP]
    SP2[K8s Container SP]
    SP3[ACM Cluster SP]

    User --> CM
    Admin --> CM
    Admin --> POL
    CM --> CDB
    CM --> PM
    PM --> POL
    PM --> SPRM
    PM --> DB
    SPRM --> AR
    SPRM -->|publish requests| MS
    MS -->|deliver requests| AG1
    MS -->|deliver requests| AG2
    AG1 --> SP1
    AG1 --> SP2
    AG2 --> SP3
    AG1 -.->|registration & heartbeat| DCM_API
    AG2 -.->|registration & heartbeat| DCM_API
    DCM_API[DCM API]
    DCM_API --> AR

    SP1 -->|status events| MS
    SP2 -->|status events| MS
    SP3 -->|status events| MS
    MS -->|status updates| SPRM
    SPRM -->|status updates| CM
  

2. Managing Policies

Policies control validation, mutation, and Agent selection for all resource requests. They are organized in a three-level hierarchy: Global (Super Admin), Tenant (Tenant Admin), and User (End User).

2.1 Create Policy

An administrator creates a policy by providing a name, type, priority, label selector, and REGO code. The Policy Manager validates uniqueness, compiles the REGO via OPA, and stores the policy metadata.

    sequenceDiagram
    actor Admin
    participant PM as Policy Manager
    participant DB as Policy DB
    participant OPA as OPA Engine

    Admin->>PM: POST /api/v1/policies<br/>{name, type, priority, label_selector, rego_code}
    PM->>PM: Validate name & priority uniqueness (at policy type level)
    alt Validation fails
        PM-->>Admin: 400 Bad Request (duplicate name or priority)
    end
    PM->>PM: Generate UUID, parse PackageName from REGO
    PM->>DB: Store policy metadata<br/>(UUID, name, package_name, label_selector, type, priority)
    PM->>OPA: Push REGO code (keyed by UUID)
    alt Compilation fails
        OPA-->>PM: Compilation error
        PM->>DB: Rollback policy record
        PM-->>Admin: 400 Bad Request (REGO compilation error)
    end
    OPA-->>PM: Compilation success
    PM-->>Admin: 201 Created {policy_id: UUID}
  

Policy payload example:

{
  "name": "restrict-region",
  "type": "Global",
  "priority": 10,
  "label_selector": { "service_type": "vm" },
  "enabled": true,
  "rego_code": "package restrict_region\n..."
}

2.2 Policy Evaluation

When a resource request arrives, the Policy Manager fetches all matching enabled policies, sorts them by level (Global > Tenant > User) then priority (ascending), and evaluates them in a chain-of-responsibility pipeline. Each policy can reject the request, apply patches (mutations), set constraints, and influence Agent selection.

    sequenceDiagram
    participant PM as Placement Manager
    participant PE as Policy Manager
    participant DB as Policy DB
    participant OPA as OPA Engine

    PM->>PE: POST /api/v1alpha1/policies:evaluateRequest<br/>{service_instance: {spec}, available_agents}

    PE->>DB: Fetch enabled policies matching request via label selector
    PE->>PE: Sort by Level (Global→Tenant→User), then Priority (asc)

    loop For each policy in sorted order
        PE->>OPA: Evaluate policy with:<br/>{spec, agent, constraints, agent_constraints}
        OPA-->>PE: {rejected, patch, constraints,<br/>selected_agent, agent_constraints}

        alt rejected == true
            PE-->>PM: 406 Not Acceptable (rejection_reason)
        end

        PE->>PE: Validate constraints<br/>(lower-level cannot unlock higher-level locks)
        alt Constraint conflict
            PE-->>PM: 409 Conflict (policy conflict error)
        end

        PE->>PE: Merge constraints into ConstraintContext
        PE->>PE: Merge agent_constraints
        PE->>PE: Validate & apply patches against constraints
        PE->>PE: Validate selected_agent against agent constraints
    end

    PE-->>PM: 200 OK {evaluated_service_instance, selected_agent, status}
  

Evaluation request (Placement Manager → Policy Manager):

{
  "service_instance": {
    "spec": {
      "service_type": "vm",
      "memory": { "size": "2GB" },
      "vcpu": { "count": 2 },
      "guest_os": { "type": "fedora-39" },
      "metadata": { "name": "fedora-vm" }
    }
  },
  "available_agents": [
    {
      "name": "agent-prod-eu-west-1",
      "environment": "prod-eu-west-1",
      "service_types": ["vm", "container"],
      "cost": "medium"
    },
    {
      "name": "agent-dev-us-east-1",
      "environment": "dev-us-east-1",
      "service_types": ["vm"],
      "cost": "low"
    }
  ]
}

Policy input (per policy, passed to OPA):

{
  "spec": {
    "service_type": "vm",
    "memory": { "size": "2GB" },
    "vcpu": { "count": 2 },
    "guest_os": { "type": "fedora-39" },
    "metadata": { "name": "fedora-vm" }
  },
  "agent": "",
  "constraints": {},
  "agent_constraints": {},
  "available_agents": [
    {
      "name": "prod-eu-agent",
      "environment": "prod-eu-west-1",
      "service_types": ["vm", "database"],
      "cost": "medium"
    }
  ],
  "exclude_agents": []
}

Policy decision format (per policy, returned by OPA):

{
  "rejected": false,
  "rejection_reason": "",
  "patch": {
    "billing_tag": "engineering",
    "region": "us-east-1"
  },
  "constraints": {
    "region": { "const": "us-east-1" },
    "vcpu": { "minimum": 2, "maximum": 8 }
  },
  "selected_agent": "agent-prod-eu-west-1",
  "agent_constraints": {
    "allow_list": ["agent-prod-eu-west-1", "agent-staging-eu-west-1"],
    "patterns": [],
    "environment_constraints": {
      "allow_list": ["prod-eu-west-1", "staging-eu-west-1"]
    }
  }
}

Evaluation response (returned to Placement Manager):

{
  "evaluated_service_instance": { "...": "final mutated spec" },
  "selected_agent": "agent-prod-eu-west-1",
  "status": "APPROVED | MODIFIED"
}

Key rules:

  • Lower-level policies cannot override constraints set by higher-level policies (e.g., a User policy cannot unlock a field locked by a Global policy).
  • A rejected: true from any policy immediately aborts evaluation (fail-fast).
  • Patches are applied cumulatively; the final payload reflects all mutations.
  • Status is APPROVED if no patches were applied, MODIFIED if the spec was mutated.

3. Managing ServiceTypes

ServiceTypes define provider-agnostic schemas for infrastructure resources. They use JSON Schema (draft 2020-12) for validation.

3.1 ServiceType Registration

ServiceTypes are defined as JSON Schemas that describe the shape of a service request. All ServiceTypes share a common structure with service_type, metadata, and optional provider_hints.

In V1, dynamic registration of ServiceType is not supported

    graph LR
    subgraph CommonFields[Common Fields]
        ST[service_type: string]
        MD[metadata: name, labels]
        PH[provider_hints: provider-specific config]
    end

    subgraph ServiceTypeSchemas[ServiceType Schemas]
        VM[VM Schema]
        CT[Container Schema]
        DBS[Database Schema]
        CL[Cluster Schema]
        STG[Storage Schema]
    end

    CommonFields --> VM
    CommonFields --> CT
    CommonFields --> DBS
    CommonFields --> CL
    CommonFields --> STG
  

3.2 Supported ServiceTypes

VM (service_type: vm)

FieldTypeRequiredDescription
vcpu.countintegeryesNumber of virtual CPUs
memory.sizestringyesMemory size (e.g., "8GB")
storage.disks[]arraynoDisks; root disk must be named "boot"
guest_os.typestringyesOS image (e.g., "rhel-9", "ubuntu-22.04")
access.ssh_public_keystringnoSSH public key for access

Container (service_type: container)

FieldTypeRequiredDescription
image.referencestringyesContainer image (e.g., "quay.io/myapp:v1.2")
resources.cpu.min/maxintegeryesCPU requests/limits
resources.memory.min/maxstringyesMemory requests/limits
process.commandarraynoEntrypoint command
process.env[]arraynoEnvironment variables
network.ports[]arraynoContainer ports

Database (service_type: database)

FieldTypeRequiredDescription
enginestringyesDatabase engine (e.g., "postgresql")
versionstringyesEngine version
resources.cpuintegeryesCPU allocation
resources.memorystringyesMemory allocation
resources.storagestringyesStorage allocation

Cluster (service_type: cluster)

FieldTypeRequiredDescription
versionstringyesKubernetes version
nodes.control_plane.countintegeryesControl plane node count (1, 3, or 5)
nodes.control_plane.cpu/memory/storagevariousyesControl plane resources
nodes.worker.countintegeryesWorker node count
nodes.worker.cpu/memory/storagevariousyesWorker node resources

Storage (service_type: storage)

FieldTypeRequiredDescription
capacitystringyesVolume size (e.g., "100Gi", "1TB")
provider_hints.kubernetes.storage_classstringnoKubernetes StorageClass name
provider_hints.kubernetes.volume_modestringnoFilesystem or Block
provider_hints.kubernetes.access_modestringnoPVC access mode (e.g., "ReadWriteOnce", "ReadWriteMany")

4. Managing CatalogItems

CatalogItems wrap ServiceType schemas with defaults, validation rules, and editability constraints. They enable administrators to create curated service offerings for end users.

4.1 Create CatalogItem

An administrator defines a CatalogItem by specifying the target ServiceType, field defaults, editability flags, and validation schemas.

    sequenceDiagram
    actor Admin
    participant CM as Catalog Manager

    Admin->>CM: POST /api/v1/catalog-items
    Note right of CM: Payload includes:<br/>- service_type reference<br/>- field definitions with:<br/>  - path (e.g., "resources.cpu")<br/>  - default value<br/>  - editable flag<br/>  - validation_schema

    CM->>CM: Validate CatalogItem schema
    CM-->>Admin: 201 Created {catalog_item_id}
  

CatalogItem example:

api_version: v1alpha1
kind: CatalogItem
metadata:
  name: production-postgres
spec:
  service_type: database
  fields:
    - path: "engine"
      default: "postgresql"
      editable: false
    - path: "version"
      editable: true
      default: "15"
      validation_schema:
        enum: ["14", "15", "16"]
    - path: "resources.cpu"
      editable: true
      default: 4
      validation_schema:
        minimum: 2
        maximum: 16
    - path: "resources.memory"
      editable: true
      default: "16GB"
    - path: "resources.storage"
      editable: true
      default: "100GB"

Storage CatalogItem example:

api_version: v1alpha1
kind: CatalogItem
metadata:
  name: standard-block-volume
spec:
  service_type: storage
  fields:
    - path: "capacity"
      editable: true
      default: "100Gi"
      validation_schema:
        type: string
        pattern: '^[0-9]+(\.[0-9]+)?(Ei|Pi|Ti|Gi|Mi|Ki|E|P|T|G|M|K)?$'
    - path: "provider_hints.kubernetes.storage_class"
      editable: false
      default: "gp3-csi"
    - path: "provider_hints.kubernetes.volume_mode"
      editable: false
      default: "Filesystem"
    - path: "provider_hints.kubernetes.access_mode"
      editable: false
      default: "ReadWriteOnce"

4.2 CatalogItem to ServiceType Translation

When a user orders an item from a CatalogItem, the system merges user input with CatalogItem defaults and validates against the field schemas, producing a ServiceType payload.

    sequenceDiagram
    actor User
    participant UI as UI / CLI
    participant CM as Catalog Manager
    participant PM as Placement Manager

    User->>UI: Select CatalogItem "production-postgres"
    UI->>UI: Render form with editable fields,<br/>defaults, and validation rules
    User->>UI: Customize editable fields<br/>(e.g., version="16", cpu=8)
    UI->>UI: Client-side validation against validation_schema
    User->>CM: POST /api/v1/catalog-item-instances<br/>{catalog_item_id, user_values}
    CM->>CM: Validate input against validation_schema
    CM->>CM: Merge defaults + user input → ServiceType payload
    Note right of CM: Result:<br/>{service_type: "database",<br/> engine: "postgresql",<br/> version: "16",<br/> resources: {cpu: 8, memory: "16GB", storage: "100GB"}}
    CM->>PM: POST /api/v1/resources<br/>{CatalogItemInstance, spec}
    PM-->>CM: 202 Accepted
    CM-->>User: Instance created (provisioning)
  

5. Service Provider & Agent Lifecycle

Service Providers register with the Environment Agent in their target environment. The Agent registers with DCM and acts as the intermediary for resource operation requests. For full details on agent behavior, see the Environment Agent enhancement.

5.1 Service Provider Registration (SP → Agent)

The Agent supports a hybrid SP model: it ships with embedded SP code for known service types (K8s Container, ACM Cluster, KubeVirt), enabled via configuration, and also accepts external (“bring your own”) SPs that register via the REST API. Only one SP — embedded or external — may serve a given service type per agent; duplicate registrations are rejected with 409 Conflict. Embedded SPs register internally at agent startup; external SPs register via POST /api/v1/providers. Registration is idempotent — re-registering with the same name updates the existing entry. External SPs periodically re-register to maintain their lease, which also ensures that after an agent restart, SPs naturally rebuild the agent’s state.

    sequenceDiagram
    participant SP as Service Provider
    participant AG as Agent
    participant DCM as DCM Control Plane
    participant DB as Database

    Note over AG: Embedded SPs registered<br/>internally at startup

    SP->>AG: POST /api/v1/providers<br/>{name, display_name, endpoint, service_type, metadata}

    alt Service type already served by another SP
        AG-->>SP: 409 Conflict<br/>{error: "service type X already served by provider Y"}
    else Name does not exist
        AG->>AG: Create new SP entry, generate provider_id
        AG-->>SP: 201 Created {id, name, status: "registered"}
    else Name exists, same provider_id
        AG->>AG: Update existing entry
        AG-->>SP: 200 OK {id, name, status: "registered"}
    else Name exists, different provider_id
        AG-->>SP: 409 Conflict
    end

    alt Service type list changed AND agent registered to DCM
        AG->>DCM: POST /api/v1/agents<br/>{name, environment, service_types, cost, topic_name}
        DCM->>DB: Update agent registration
        DCM-->>AG: 200 OK
    end

    Note over SP,AG: SP periodically re-registers<br/>to maintain lease
  

Registration payload example (SP → Agent):

{
  "name": "kubevirt-sp",
  "display_name": "KubeVirt Service Provider",
  "endpoint": "https://sp1.example.com/api/v1/vm",
  "service_type": "vm",
  "metadata": {
    "region": "us-east-1",
    "resources": {
      "total_cpu": 200,
      "total_memory": "1TB",
      "total_storage": "2TB"
    }
  }
}

5.2 Agent Registration (Agent → DCM)

The Agent registers with DCM after creating its messaging topics and after at least one SP (embedded or external) is registered and healthy. Registration is idempotent — the agent name is the natural key. On restart, the agent re-registers; DCM resets the heartbeat tracker. For full registration details, see the Environment Agent enhancement.

    sequenceDiagram
    autonumber
    participant AG as Agent
    participant MS as Messaging System
    participant DCM as DCM Control Plane
    participant DB as Database

    AG->>MS: Create topics (main + retry)
    Note over AG: Wait for at least 1 SP<br/>(embedded or external) to register<br/>and be healthy

    AG->>DCM: POST /api/v1/agents<br/>{name, environment, service_types,<br/>resources_available, cost, topic_name}
    DCM->>DB: Store agent registration
    DCM-->>AG: 201 Created {agent_id}
  

Registration payload (Agent → DCM):

{
  "name": "agent-prod-eu-west-1",
  "environment": "prod-eu-west-1",
  "service_types": ["vm", "container"],
  "resources_available": {
    "total_cpu": 200,
    "total_memory": "1TB",
    "total_storage": "2TB"
  },
  "cost": "medium",
  "topic_name": "dcm.agents.agent-prod-eu-west-1"
}

5.3 Health Monitoring

5.3.1 SP Health (Agent → SP)

The Agent monitors each registered SP’s health using a three-state model. The monitoring mechanism differs by SP type: embedded SPs are checked in-process (no network call), while external SPs are checked by polling their /health endpoint at a configurable interval.

Health State Diagram
    stateDiagram-v2
    [*] --> Ready: Registered with Agent

    Ready --> FailureCount: Failed
    FailureCount --> Ready: OK (reset)
    FailureCount --> Unavailable: Threshold reached

    Ready --> Unhealthy: status: unhealthy
    Unhealthy --> Ready: status: healthy
    Unhealthy --> Unavailable: Timeout/error threshold

    Unavailable --> Ready: OK (recover)
  

Three health states:

  • Ready: SP is healthy and eligible for routing.
  • Unhealthy: SP is reachable but reports its backing provider is down. The Agent keeps the service type in its advertised list but stops routing requests to this SP; incoming requests are held in the retry topic until the SP recovers or becomes Unavailable.
  • Unavailable: SP is unreachable after exceeding the failure threshold. The Agent removes the service type from its advertised list and updates DCM.
Health Check Sequence
    sequenceDiagram
    participant AG as Agent
    participant SP as External SP

    Note over AG: Embedded SPs: health<br/>checked in-process<br/>(no network call)

    loop Every {health_check_interval} seconds (external SPs only)
        AG->>SP: GET /health
        alt 200 OK, status: healthy
            SP-->>AG: {status: "healthy"}
            AG->>AG: Reset failure counter, mark Ready
        else 200 OK, status: unhealthy
            SP-->>AG: {status: "unhealthy"}
            AG->>AG: Mark Unhealthy<br/>Stop routing, hold requests
        else Timeout or error
            SP-->>AG: Error / Timeout
            AG->>AG: Increment failure counter
            alt Failures >= threshold
                AG->>AG: Mark Unavailable<br/>Remove service type, update DCM
            end
        end
    end
  

5.3.2 Agent Health (Agent → DCM heartbeats)

The Agent reports its own liveness to DCM via periodic REST heartbeats. DCM tracks the last heartbeat timestamp for each agent and marks the agent as Unavailable if no heartbeat is received within a configurable threshold.

    sequenceDiagram
    participant AG as Agent
    participant DCM as DCM Control Plane

    loop Every {heartbeat_interval} seconds
        AG->>DCM: PUT /api/v1/agents/{agent_id}/heartbeat<br/>{timestamp, consumer_lag}
        DCM->>DCM: Update heartbeat, check lag
        DCM-->>AG: 200 OK
    end

    Note over DCM: No heartbeat within threshold
    DCM->>DCM: Mark agent Unavailable
  

5.3.3 Consumer Lag Monitoring

The Agent self-reports its consumer lag in each heartbeat. If the lag exceeds consumer_lag_threshold, DCM marks the agent as Congested and stops routing new requests to it. When the lag drops below the threshold, the Congested state is cleared.

Agent Health State Diagram
    stateDiagram-v2
    [*] --> Ready: Agent registers
    Ready --> Congested: lag >= threshold
    Congested --> Ready: lag < threshold
    Ready --> Unavailable: Heartbeat timeout
    Congested --> Unavailable: Heartbeat timeout
    Unavailable --> Ready: Agent re-registers
  

5.4 Service Provider Status Reporting

Note: Status reporting is not impacted by the Agent layer. SPs publish status CloudEvents directly to the Messaging System. The Agent is not in the status-reporting path.

Service Providers report instance status changes to DCM via CloudEvents published to a messaging system (NATS). This decoupled approach supports multiple consumers (billing, auditing, etc.) and scales independently.

    sequenceDiagram
    participant Platform as Underlying Platform<br/>(K8s, KubeVirt, ACM)
    participant SP as Service Provider
    participant MS as Messaging System (NATS)
    participant DCM as DCM Core Service
    participant DB as Status DB

    Platform->>SP: State change event<br/>(via informer watch or polling)
    SP->>SP: Map platform status → DCM status
    SP->>SP: Build CloudEvent
    SP->>MS: Publish to NATS subject<br/>dcm.{service_type}<br/>

    MS->>DCM: Deliver event
    DCM->>DCM: Validate CloudEvent schema
    alt Valid
        DCM->>DB: UPSERT instance status
    else Invalid
        DCM->>DCM: Log error, discard
    end
  

Status enums by ServiceType:

VMContainerClusterStorage
PROVISIONINGPENDINGPROGRESSINGPROVISIONING
RUNNINGRUNNINGACTIVERUNNING
STOPPEDSUCCEEDEDFAILED
PAUSEDFAILEDDEGRADED
FAILEDUNKNOWNDELETEDFAILED
DELETINGUNAVAILABLEDELETING
DELETEDDELETINGDELETED

5.5 Agent Lifecycle

This section provides a brief overview of the agent lifecycle. For full details, see the Environment Agent enhancement.

Startup:

  1. Agent registers its configured embedded SPs internally (K8s Container, ACM Cluster, KubeVirt — each if enabled in config)
  2. Agent creates messaging topics (main topic + retry topic)
  3. Agent waits for at least one SP (embedded or external) to be registered and healthy
  4. Agent registers with DCM via POST /api/v1/agents
  5. Agent begins periodic heartbeats and SP health checking

Restart:

  1. Agent re-registers with DCM (idempotent; DCM resets heartbeat tracker)
  2. Embedded SPs register internally at startup; external SPs naturally re-register via periodic lease renewal, rebuilding agent state
  3. Unconsumed messages on both main and retry topics survive (messaging system persistence)
  4. Agent resumes consuming from both topics once fully initialized

6. CatalogItemInstance Creation (End-to-End)

This is the primary user flow: creating an infrastructure resource from a CatalogItem. The request flows through the Catalog Manager, Placement Manager (with policy evaluation), SP Resource Manager (which publishes to the messaging system), the Environment Agent, and finally to the selected Service Provider.

6.1 Full Creation Flow

    sequenceDiagram
    actor User
    participant CM as Catalog Manager
    participant PM as Placement Manager
    participant DB as Placement DB
    participant PE as Policy Manager
    participant SPRM as SP Resource Manager
    participant AR as Agent Registry
    participant MS as Messaging System
    participant AG as Agent
    participant SP as Service Provider

    User->>CM: Request CatalogItemInstance<br/>(select CatalogItem + customize fields)
    CM->>CM: Validate input, merge with defaults
    CM->>PM: POST /api/v1/resources<br/>{CatalogItemInstance: UUID, spec}

    PM->>DB: Store original request (intent)

    PM->>AR: Fetch available agents<br/>(healthy, not Congested, matching service_type)
    AR-->>PM: available_agents list

    PM->>PE: POST /api/v1alpha1/policies:evaluateRequest<br/>{service_instance: {spec}, available_agents}
    PE->>PE: Evaluate policy chain<br/>(validate, mutate, select Agent)

    alt Policy rejects
        PE-->>PM: 406 Not Acceptable
        PM->>DB: Delete intent record
        PM-->>CM: Error (policy rejected)
        CM-->>User: Request denied
    end

    PE-->>PM: 200 OK<br/>{evaluated_service_instance, selected_agent, status}
    PM->>DB: Store validated request with agent_name

    PM->>SPRM: POST /api/v1/service-type-instances<br/>{agent_name, service_type, spec}

    SPRM->>AR: Lookup agent, get topic_name
    alt Agent not found or unhealthy/Congested
        SPRM-->>PM: Error (404/503)
        PM->>DB: Delete records
        PM-->>CM: Error
        CM-->>User: Agent unavailable
    end

    SPRM->>MS: PUBLISH CloudEvent<br/>topic: {topic_name}<br/>{resource_id, service_type, spec}
    SPRM->>DB: Create instance record
    SPRM-->>PM: 202 Accepted {instance_id, agent_name, status: PENDING}
    PM-->>CM: 201 Created
    CM-->>User: Instance created (PENDING)

    Note over MS,AG: Async processing
    MS->>AG: Deliver creation request
    AG->>AG: Validate service type, select SP
    AG->>SP: POST {sp_endpoint}/api/v1/{service_type}<br/>{spec}
    SP-->>AG: {instance_id, status: PROVISIONING}
    AG->>MS: PUBLISH CloudEvent<br/>topic: dcm.agents.responses<br/>{resource_id, agent_name, topic_name,<br/>status: PROVISIONING}
    MS->>SPRM: Deliver response
    SPRM->>DB: Update instance: PROVISIONING

    opt Agent queues request (SP Unhealthy)
        AG->>MS: PUBLISH CloudEvent<br/>topic: dcm.agents.responses<br/>{resource_id, status: QUEUED}
        MS->>SPRM: Deliver QUEUED response
        SPRM->>SPRM: Update instance: QUEUED
        SPRM->>PM: Notify: instance QUEUED

        Note over PM: Start queued_request_timeout
        alt Timeout or timeout = 0
            PM->>SPRM: DELETE instance
            PM->>PE: Re-evaluate excluding agent
            Note over PM: Route to alternative agent<br/>or return error if none available
        end
    end

    Note over SP,MS: Status reporting (unchanged)
    SP->>MS: Publish status CloudEvents
    MS->>SPRM: Deliver status updates
  

When the SP for the requested service type on the agent is Unhealthy, the Agent holds the request in its retry topic and responds with a QUEUED CloudEvent. DCM records the QUEUED status. If the SP recovers, the Agent processes the held request. If the SP becomes Unavailable, the Agent rejects the held request with an error CloudEvent. The Placement Manager handles the QUEUED status via a queued_request_timeout timer (see 6.2 Placement Manager Flow).

6.2 Placement Manager Flow

The Placement Manager is the central orchestrator. It preserves the user’s original intent, fetches available agents, delegates policy evaluation, and coordinates with the SP Resource Manager.

    flowchart TD
    A[Receive request from Catalog Manager] --> B[Store original request in Placement DB]
    B --> C[Fetch available agents from Agent Registry]
    C --> D[Send to Policy Manager for evaluation<br/>with available_agents]
    D --> E{Policy approved?}
    E -->|No| F[Delete intent record]
    F --> G[Return error to Catalog Manager]
    E -->|Yes| H[Store validated request with agent_name]
    H --> I[Forward to SP Resource Manager<br/>with agent_name, service_type, spec]
    I --> J{SPRM response?}
    J -->|Error| K[Delete records from Placement DB]
    K --> G
    J -->|202 Accepted| L[Return 201 Created<br/>to Catalog Manager]
    J -->|QUEUED| M[Start queued_request_timeout timer]
    M --> N{Timeout?}
    N -->|Yes| O[Send DELETE to SPRM<br/>Re-evaluate excluding agent]
    O --> P{Alternative agent?}
    P -->|Yes| I
    P -->|No| K
  

Request payload (Catalog Manager → Placement Manager):

{
  "CatalogItemInstance": "4baa35eb-e70d-4d37-867d-0f4efa21d05c",
  "spec": {
    "service_type": "vm",
    "memory": { "size": "2GB" },
    "vcpu": { "count": 2 },
    "guest_os": { "type": "fedora-39" },
    "access": { "ssh_public_key": "ssh-ed25519 ..." },
    "metadata": { "name": "fedora-vm" }
  }
}

Response payload (Placement Manager → Catalog Manager):

{
  "CatalogItemInstanceId": "f3645f8f-82c1-4efb-888f-318c0ac81a08",
  "resource_name": "fedora-vm",
  "agent_name": "agent-prod-eu-west-1",
  "id": "08aa81d1-a0d2-4d5f-a4df-b80addf07781"
}

6.3 SP Resource Manager Flow

The SP Resource Manager handles Agent lookup and publishes creation requests as CloudEvents to the agent’s messaging topic. It no longer calls SP REST endpoints directly.

    flowchart TD
    A[Receive request from Placement Manager<br/>agent_name + service_type + spec] --> B[Query Agent Registry<br/>by agent_name]
    B --> C{Agent found?}
    C -->|No| D[Return 404 Not Found]
    C -->|Yes| E{Agent healthy<br/>and not Congested?}
    E -->|No| F[Return 503 Service Unavailable]
    E -->|Yes| G[Publish CloudEvent to agent topic<br/>via Messaging System]
    G --> H[Create instance record in DB]
    H --> I[Return 202 Accepted<br/>instance_id, agent_name, status: PENDING]
  

6.4 Service Provider Instance Creation

Each Service Provider translates the provider-agnostic ServiceType spec into platform-native resources.

Note: Each service type is served by exactly one SP (embedded or external) per agent — there is no SP selection strategy. The Agent forwards the request to the SP via an in-process call (for embedded SPs) or via REST (for external SPs). The SP’s internal behavior is unchanged.

    flowchart LR
    subgraph KubeVirtSP[KubeVirt SP]
        A1[Receive VM spec] --> A2[Create VirtualMachine CR]
        A2 --> A3[Return instance_id - PROVISIONING]
    end

    subgraph K8sContainerSP[K8s Container SP]
        B1[Receive Container spec] --> B2[Create Deployment and Service]
        B2 --> B3[Return request_id - PENDING]
    end

    subgraph ACMClusterSP[ACM Cluster SP]
        C1[Receive Cluster spec] --> C2[Create HostedCluster and NodePool]
        C2 --> C3[Return request_id - PENDING]
    end

    subgraph K8sStorageSP[K8s Storage SP]
        D1[Receive Storage spec] --> D2[Create PersistentVolumeClaim]
        D2 --> D3[Return request_id - PROVISIONING]
    end
  

6.5 Continuous Status Reporting

Note: Status reporting is not impacted by the Agent layer. SPs publish status CloudEvents directly to the Messaging System, bypassing the Agent. This path is unchanged from the pre-agent architecture.

After instance creation, Service Providers continuously monitor the underlying platform and report status changes via CloudEvents.

    flowchart TD
    subgraph Service Provider
        A[Platform event detected<br/>via Informer watch or polling]
        A --> B[Map platform status<br/>to DCM status enum]
        B --> C[Build CloudEvent v1.0]
        C --> D["Publish to NATS<br/>dcm.{service_type}"]
    end

    subgraph DCM Core
        D --> E[Receive CloudEvent]
        E --> F{Valid schema?}
        F -->|Yes| G[UPSERT instance status in DB]
        F -->|No| H[Log error, discard]
    end

    subgraph Monitoring Approaches
        I[Event-Driven Streaming<br/>Preferred - K8s Informers]
        J[Polling<br/>Fallback - Legacy APIs]
    end
  

Platform status mapping examples:

    graph LR
    subgraph KubeVirt VMI Phase
        VP1[Pending/Scheduling/Scheduled]
        VP2[Running]
        VP3[Succeeded]
        VP4[Failed/Unknown]
        VP5[Not Found]
    end

    subgraph DCM VM Status
        DS1[PROVISIONING]
        DS2[RUNNING]
        DS3[STOPPING]
        DS4[FAILED]
        DS5[DELETED]
    end

    VP1 --> DS1
    VP2 --> DS2
    VP3 --> DS3
    VP4 --> DS4
    VP5 --> DS5
  
    graph LR
    subgraph K8s Pod Phase
        KP1[Pending/ContainerCreating]
        KP2[Running]
        KP3[Succeeded]
        KP4[Failed/CrashLoopBackOff]
        KP5[Unknown - node lost]
    end

    subgraph DCM Container Status
        CS1[PENDING]
        CS2[RUNNING]
        CS3[SUCCEEDED]
        CS4[FAILED]
        CS5[UNKNOWN]
    end

    KP1 --> CS1
    KP2 --> CS2
    KP3 --> CS3
    KP4 --> CS4
    KP5 --> CS5
  
    graph LR
    subgraph HostedCluster Conditions
        HC1[Progressing=Unknown]
        HC2[Progressing=True, Available=False]
        HC3[Available=True, Progressing=False]
        HC4[Degraded=True]
        HC5[Not Found]
    end

    subgraph DCM Cluster Status
        DC1[PENDING]
        DC2[PROVISIONING]
        DC3[READY]
        DC4[FAILED]
        DC5[DELETED]
    end

    HC1 --> DC1
    HC2 --> DC2
    HC3 --> DC3
    HC4 --> DC4
    HC5 --> DC5
  
    graph LR
    subgraph K8s PVC Phase
        PVC1[Pending]
        PVC2[Bound - resizing]
        PVC3[Bound]
        PVC4[Lost]
        PVC5[deletionTimestamp set]
        PVC6[Not Found]
    end

    subgraph DCM Storage Status
        SS1[PROVISIONING]
        SS2[PROVISIONING]
        SS3[RUNNING]
        SS4[FAILED]
        SS5[DELETING]
        SS6[DELETED]
    end

    PVC1 --> SS1
    PVC2 --> SS2
    PVC3 --> SS3
    PVC4 --> SS4
    PVC5 --> SS5
    PVC6 --> SS6
  

6.6 Deletion Flow

The deletion flow follows the same architecture as creation: the request is published as a CloudEvent to the agent’s messaging topic, and the Agent routes it to the appropriate SP.

    sequenceDiagram
    actor User
    participant CM as Catalog Manager
    participant PM as Placement Manager
    participant DB as Placement DB
    participant SPRM as SP Resource Manager
    participant MS as Messaging System
    participant AG as Agent
    participant SP as Service Provider

    User->>CM: Delete CatalogItemInstance
    CM->>PM: DELETE /api/v1/resources/{resource_id}
    PM->>DB: Lookup resource (agent_name, service_type, instance_id)

    PM->>SPRM: DELETE /api/v1/service-type-instances/{instance_id}
    SPRM->>MS: PUBLISH CloudEvent<br/>topic: {topic_name}<br/>type: dcm.request.delete<br/>{resource_id, service_type}
    SPRM-->>PM: 202 Accepted

    MS->>AG: Deliver deletion request
    AG->>SP: DELETE {sp_endpoint}/api/v1/{service_type}/{resource_id}
    SP-->>AG: {status: DELETING}
    AG->>MS: PUBLISH CloudEvent<br/>topic: dcm.agents.responses<br/>{resource_id, agent_name, topic_name,<br/>status: DELETING}

    opt Agent queues request (SP Unhealthy)
        AG->>MS: PUBLISH CloudEvent<br/>topic: dcm.agents.responses<br/>{resource_id, status: QUEUED}
        MS->>SPRM: Deliver QUEUED response
        SPRM->>SPRM: Update instance: QUEUED
        SPRM->>PM: Notify: deletion QUEUED

        Note over PM: Resource stays DELETING.<br/>Deletion cannot be re-routed.<br/>Agent holds the request in its<br/>retry topic for automatic resolution.

        alt SP recovers — Agent processes held deletion
            AG->>SP: DELETE {sp_endpoint}/api/v1/{service_type}/{resource_id}
            SP-->>AG: {status: DELETING}
            AG->>MS: PUBLISH CloudEvent<br/>{resource_id, status: DELETING}
        else SP becomes Unavailable — Agent rejects
            AG->>MS: PUBLISH CloudEvent<br/>{resource_id, error: "SP unavailable"}
            MS->>SPRM: Deliver error
            Note over SPRM: Enqueue in cleanup queue<br/>for deferred retry.<br/>Resource stays DELETING.
        end
    end

    Note over SP: SP manages deletion<br/>and reports final status
    SP->>MS: CloudEvent {status: DELETED}
    MS->>SPRM: Status update
    SPRM->>DB: Update status: DELETED