Metadata-Version: 2.4
Name: datadog-kueue
Version: 1.1.0
Summary: The Kueue check
Project-URL: Source, https://github.com/DataDog/integrations-core
Author-email: Datadog <packages@datadoghq.com>
License-Expression: BSD-3-Clause
Keywords: datadog,datadog agent,datadog check,kueue
Classifier: Private :: Do Not Upload
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: System Administrators
Classifier: License :: OSI Approved :: BSD License
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: System :: Monitoring
Requires-Python: >=3.12
Requires-Dist: datadog-checks-base>=37.33.0
Provides-Extra: deps
Requires-Dist: kubernetes==35.0.0; extra == 'deps'
Description-Content-Type: text/markdown

# Agent Check: Kueue

## Overview

This check monitors Kueue through the Datadog Agent.

Kueue is a Kubernetes workload queueing system that allows you to manage and schedule workloads on your Kubernetes cluster. It provides a way to prioritize and manage workloads, and to ensure that workloads are scheduled in a fair and efficient manner. This integration collects metrics from the Kueue controller manager and Kueue API server to help you monitor the health and performance of your Kueue cluster.

## Setup

Follow the instructions below to install and configure this check for an Agent running on a host. For containerized environments, see the [Autodiscovery Integration Templates][3].

### Installation

The Kueue check is included in the [Datadog Agent][2] package.
No additional installation is required on your server.

### Configuration

Kueue is a cluster-level service. Configure this integration as a Cluster Agent cluster check so only one Agent instance scrapes the Kueue metrics endpoint.

1. To collect optional ClusterQueue resource metrics, such as `kueue.cluster_queue.resource_usage.gpu`, configure Kueue with `metrics.enableClusterQueueResources: true` and restart the Kueue controller manager.

2. Provide a [cluster check configuration][10] to the Cluster Agent. For file or ConfigMap based configuration, set `cluster_check: true` in the instance:

   ```yaml
   clusterAgent:
     confd:
       kueue.yaml: |-
         cluster_check: true
         init_config:
         instances:
         - openmetrics_endpoint: http://kueue-controller-manager-metrics-service.kueue-system.svc:8080/metrics
   ```

   Kueue Workload lifecycle events are collected by default. The Agent running the check needs `get` and `list`
   permissions on the `workloads` resource in the `kueue.x-k8s.io` API group. Set `collect_workload_events: false` to
   disable event collection.

3. Alternatively, annotate the Kueue metrics service with Autodiscovery cluster check annotations:

   ```yaml
   ad.datadoghq.com/endpoints.checks: |
     {
       "kueue": {
         "instances": [
           {
             "openmetrics_endpoint": "http://%%host%%:%%port%%/metrics"
           }
         ]
       }
     }
   ```

See the [sample kueue.d/conf.yaml][4] for all available configuration options.

### Log collection

The Kueue controller manager writes logs to its container output, which Kubernetes captures as container logs. Collecting logs is disabled by default in the Datadog Agent. To enable it, see [Kubernetes Log Collection][12]. Logs are collected by the node Agent running on the node that hosts the Kueue controller manager, not by the Cluster Agent that runs this cluster check.

After log collection has been enabled, set the Kueue log configuration as an Autodiscovery annotation on the controller manager's pod template. This allows it to persist despite pod restarts. Add it under `spec.template.metadata.annotations` of the `kueue-controller-manager` deployment, or set `controllerManager.manager.podAnnotations` if you install Kueue with the Helm chart:

```yaml
ad.datadoghq.com/manager.logs: |
  [
    {
      "source": "kueue",
      "service": "<SERVICE>"
    }
  ]
```

This annotation targets the container named `manager`, which is the container name used by both the Kueue release manifests and the Helm chart. Replace `manager` with the name (`.spec.containers[i].name`) of your Kueue container if you use a different name.

### Validation

[Run the Cluster Agent's `clusterchecks` subcommand][11] and look for `kueue` under the Checks section.

## Data Collected

### Metrics

See [metadata.csv][7] for a list of metrics provided by this integration.

### Events

By default, the Kueue integration polls the Kueue Workload custom resources and sends Datadog events for lifecycle
transitions:

- `kueue.workload.created`: a Workload appears after the check has initialized its state.
- `kueue.workload.quota_reserved`: the `QuotaReserved` condition becomes `True`.
- `kueue.workload.admitted`: the `Admitted` condition becomes `True`.
- `kueue.workload.running`: the `PodsReady` condition becomes `True`.
- `kueue.workload.evicted`: the `Evicted` condition becomes `True`.
- `kueue.workload.finished`: the `Finished` condition becomes `True`.

Events are tagged with the Workload namespace, name, UID, LocalQueue, transition, priority, and ClusterQueue when
available. Eviction events also include the eviction reason, and preemption events include the Kueue preemption reason
when available.

The first collection run seeds the Workload state and does not emit events for already existing transitions.

## Troubleshooting

Need help? Contact [Datadog support][8].


[2]: https://app.datadoghq.com/account/settings/agent/latest
[3]: https://docs.datadoghq.com/containers/kubernetes/integrations/
[4]: https://github.com/DataDog/integrations-core/blob/master/kueue/datadog_checks/kueue/data/conf.yaml.example
[5]: https://docs.datadoghq.com/agent/configuration/agent-commands/#start-stop-and-restart-the-agent
[6]: https://docs.datadoghq.com/agent/configuration/agent-commands/#agent-status-and-information
[7]: https://github.com/DataDog/integrations-core/blob/master/kueue/metadata.csv
[8]: https://docs.datadoghq.com/help/
[10]: https://docs.datadoghq.com/containers/cluster_agent/clusterchecks/?tab=helm#configuration-from-configuration-files
[11]: https://docs.datadoghq.com/containers/troubleshooting/cluster-and-endpoint-checks/#dispatching-logic-in-the-cluster-agent
[12]: https://docs.datadoghq.com/containers/kubernetes/log/
