> For the complete documentation index, see [llms.txt](https://docs.datumo.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.datumo.com/documentation/documentation-en/redteaming/how-to-use/2.-run-auto-redteaming.md).

# 2. Run Auto Redteaming

## 1. Run Red Teaming Evaluation (Run Attack Set)

> 💡 Overview
>
> When an Attack Set is created, an automatic red teaming evaluation is run, and various red teaming strategies are automatically applied to verify the model’s safety.

***

### Screen Layout <a href="#undefined" id="undefined"></a>

Auto Red Teaming is **Task → Attack Set → Result** structured as.

| Screen                        | Description                                              | Main Actions                                    |
| ----------------------------- | -------------------------------------------------------- | ----------------------------------------------- |
| Task List                     | Display the created Task list and overall status         | Create Task / Select Task                       |
| Task Details - Dashboard Tab  | Summary of overall results by Task                       | Compare results by model, view detailed results |
| Task Details - Attack Set Tab | List of Attack Sets included in the Task                 | Check Attack Set / Auto Red Teaming Run status  |
| Attack Set Details            | Execution status and results of an individual Attack Set | Check Auto Red Teaming results                  |

***

### 1. Create Task <a href="#step-1-task" id="step-1-task"></a>

A Task is the **basic unit**of the management page for the evaluation, and also a **container** that groups multiple Attack Sets.

<figure><img src="/files/2wSw3uhWPcWoArNyaL06" alt=""><figcaption></figcaption></figure>

**① Click + New Task**

At the top right of the Task list, **+ New Task** Click the button.

**② Enter Task information**

| Item                   | Description                               |
| ---------------------- | ----------------------------------------- |
| Task Name (required)   | Task name (up to 255 characters)          |
| Description (optional) | Task description (up to 1,000 characters) |

**③ Complete**

**Complete** 버튼을 클릭하면 Task가 생성되고 목록으로 이동합니다.

***

**③ Complete**

**Complete** When you click the button, the Task is created and you are moved to the list.

***

### 2. Create Attack Set <a href="#step-2-attack-set" id="step-2-attack-set"></a>

An Attack Set is the execution unit where the actual red teaming evaluation is performed. When you add an Attack Set from the Task details screen, the evaluation runs automatically after setup is complete.

**① Enter Task details**

Click the desired Task row in the Task list to move to the Task details screen.

<figure><img src="/files/uSpQbSWkRYzUdDMI7qDl" alt=""><figcaption></figcaption></figure>

**② Click + Add Attack Set**

In the Attack Set tab, **+ Add Attack Set** Click the button.

<figure><img src="/files/S0te17ohBsnI1r1xqQlI" alt=""><figcaption></figcaption></figure>

<br>

***

The Attack Set creation modal is structured as follows.

* Left area: Select the Benchmark Dataset to use for evaluation
* Right area: Enter evaluation settings for running the Attack Set

**③ Select Dataset (left)**

Select the Benchmark Dataset to use for red teaming evaluation.

* A Dataset is a set of Seeds organized by Risk Taxonomy.
* You can filter the Dataset list through search.

<figure><img src="/files/AiNVntgpPgmG3s5n08dn" alt=""><figcaption></figcaption></figure>

**④ Enter settings (right)**

<table><thead><tr><th width="78.32421875">Step</th><th>Item</th><th>Description</th></tr></thead><tbody><tr><td>1</td><td>Attack Set Name / Description</td><td>Name (required), description (optional)</td></tr><tr><td>2</td><td>Target Model/Agent</td><td>Model to evaluate (multiple selection available)</td></tr><tr><td>3</td><td>Max Red Teaming Runs</td><td>Maximum number of attacks per seed (default 20, max 50)</td></tr><tr><td>4</td><td><p>Evaluation Sampling M</p><p>ethod</p></td><td>(When sampling is in progress) Select the dataset sampling method</td></tr></tbody></table>

<figure><img src="/files/RTK9uWEzKO4XjOZBWaTb" alt=""><figcaption></figcaption></figure>

<details>

<summary>📂 Sampling Method Details</summary>

</details>

> 💡 Information
>
> The number of Seeds included in a Benchmark Dataset may differ by taxonomy.\
> The Evaluation Sampling Method is an option for selecting how such differences are reflected in the evaluation.

> 💡 Information
>
> If multiple models are selected,\
> Attack Sets that share the same Dataset and settings are created separately for each model,\
> allowing direct comparison of results.

***

### 3. Run Evaluation <a href="#step-3" id="step-3"></a>

Once the Attack Set setup is complete, you can run the red teaming evaluation.

**① Click Complete**

After setup is complete **Complete** when you click the button, before the evaluation is run, **the final confirmation modal**appears.

<figure><img src="/files/NprrPPfqAbCWjAkF23nH" alt=""><figcaption></figcaption></figure>

> ⚠️ Check before running
>
> * Once the evaluation starts, it cannot be paused or stopped.
> * The execution time may vary depending on the size of the selected Dataset and the settings.

**② Click Proceed**

Clicking Proceed starts the red teaming evaluation immediately.\
When the evaluation starts, the Attack Set status becomes **In Progress (Red teaming in progress)**&#x61;nd is displayed as,\
and the evaluation continues to run in the background even if you leave the page.

<figure><img src="/files/Y5GVZDOjw1NoaXfrbrZM" alt=""><figcaption></figcaption></figure>

**③ Check execution status**

The evaluation progress can be checked in the Attack Set list and on the details screen.

| Status      | Description                  |
| ----------- | ---------------------------- |
| Waiting     | Waiting                      |
| In Progress | In progress (progress shown) |
| Done        | Done                         |
| Error       | Error                        |

<figure><img src="/files/7G5bzjpfvNyD7cV8Xa1o" alt=""><figcaption></figcaption></figure>

> 💡 Background execution
>
> Even if you navigate away from the page or close the window after starting the evaluation, the running evaluation will not stop.

***

### 4. Management Features <a href="#step-4" id="step-4"></a>

In Auto Red Teaming, to ensure reproducibility and integrity of evaluation results,\
limited management features are provided for Tasks and Attack Sets.

#### 1. Task Management <a href="#id-1-task" id="id-1-task"></a>

A Task is a higher-level unit that groups multiple Attack Sets, and editing or deleting at the Task level is restricted depending on the status of the child Attack Sets.

**① Edit Task**

On the Task list screen, you can edit the Task Name / Description by clicking the Edit button.

* Can be edited regardless of whether an evaluation is running
* Does not affect evaluation results

**② Delete Task**

> ⚠️ Deletion conditions
>
> Task deletion is performed on the Task list screen.
>
> * A Task can be deleted only when all Attack Sets included in the Task are in Done status.
> * When deleted, all Attack Sets included in that Task and their evaluation result data are deleted together.
> * Deleted data cannot be restored

💡 If there are any running or unfinished evaluations, Task deletion is restricted to protect the integrity of the results.

<figure><img src="/files/woUXMj6wWnorXRgKbmyI" alt=""><figcaption></figcaption></figure>

***

#### 2. Attack Set Management <a href="#id-2-attack-set" id="id-2-attack-set"></a>

An Attack Set is the execution unit where the actual red teaming evaluation is performed, and changes to evaluation conditions are not allowed.

| Item                 | Manageable or not | Description                                |
| -------------------- | ----------------- | ------------------------------------------ |
| Name / Description   | Editable          | For identification and management purposes |
| Dataset              | Not editable      | Ensures evaluation reproducibility         |
| Target Model / Agent | Not editable      | Ensures comparison reliability             |
| Sampling Method      | Not editable      | Maintains result consistency               |
| Delete               | `Done` 상태에서만 가능   | Protection during execution                |

> 💡 Attack Set deletion is performed from the individual row in Task Details > Attack Set tab.

***

### 🔗 Related Documents

* R-1. Benchmark Set Management — How to check the composition of Seeds and Risk Taxonomy for attack simulation
* R-2. Auto Red Teaming — How to configure an Attack Set, select a Target Model, and run automatic red teaming
* R-3. Dashboard — How to visualize and analyze safety evaluation results such as ASR and Score
