> For the complete documentation index, see [llms.txt](https://academy.gooey.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://academy.gooey.ai/ai-for-impact/module-8/how-to-use-bulk-runner.md).

# How to use bulk runner?

{% embed url="<https://youtu.be/2_k3Zg4Z1Rg>" %}

**1. Prepare Your Test Questions and Golden Answers**

* Create a spreadsheet with your test questions and golden answers.
* Your sheet should have columns for:&#x20;
  * question
  * golden answer
  * citation (if needed)
  * audio file as a Google Drive link if needed

<figure><img src="/files/As9bfCuDMNftNJdY7Li2" alt=""><figcaption></figcaption></figure>

**2. Create or Duplicate Your AI Agent**

* You need a AI Agent for each model or prompt you want to test.
* To create a new AI Agent for a different model (for example, to test Gemini 2.5 Pro vs. GPT 4.1):
  * Go to your existing AI Agent, click "Update," then "Save as new" to duplicate it.
  * Choose the new model (such as Gemini 2.5).
  * Update the name (for example, "Gemini 2.5").
  * Click "Save".

**3. Set Up the Bulk Run**

* Go to [gooey.ai/bulk](https://gooey.ai/bulk).
* Link your spreadsheet containing the test questions and golden answers:
  * In the "Input data spreadsheet" section, click "Link" and paste your spreadsheet URL.
  * Click "Import."
* Once imported, check that your questions and golden answers have loaded correctly.

<figure><img src="/files/3y4qqIBpllG2mCFB9D2G" alt=""><figcaption></figcaption></figure>

**4. Add Your AI Agents as Workflows**

* In the bulk runner, click "Add workflow."
* Start typing the name of your AI Agent (for example, "marketing\_gooey\_support\_bot") and select it.
* Add each AI Agent you want to compare (for example, one for GPT 4.0, one for Gemini 2.5 Pro).

<figure><img src="/files/8AvWt8Jk4xxhsxjoeE6W" alt=""><figcaption></figcaption></figure>

**5. Configure the Input and Output Columns**

* Go to "Show all columns."
* Set "Input prompt" to your question column (e.g., "question").
* Make sure "Output text," "Run URL," and "Runtime" are checked. They help you with results and debugging.

<figure><img src="/files/u4hOd9aZy2SsGSY2zgGG" alt=""><figcaption></figcaption></figure>

**6. Enable Evaluation Workflow**

* In the "Evaluation workflows" section, enable "Copilot evaluator."
* This will compare each model's output to your golden answer and score them.

<figure><img src="/files/24ZKsfTfbC9UbiLr9JY7" alt=""><figcaption></figcaption></figure>

{% hint style="info" %}
If you only want to run the bulk runner without evaluation, you can delete the evaluator.
{% endhint %}

**7. Start the Bulk Run**

* Click "Run."
* Gooey.AI will process each question through every selected Copilot/model.
* For each question and Copilot, you get the generated answer, run URL, runtime, and more.

**8. Review and Compare Results**

* In the results sheet:
  * Each row shows the question, the answer from each Copilot, the runtime, and the run URL.
  * At the end, you will see the evaluation scores for each model.
  * The system identifies which model performed the best for each question and overall.

**9. Analyze Performance**

* Look at the evaluation scores (for example, 80%, 100%, 60%).
* Higher scores mean answers closer to the expert-provided golden answer.
* If a new model scores lower, review the answers and ratings to find areas for improvement.

**10. Repeat or Refine**

* You can rerun the evaluation after adjusting prompts, models, or questions.
* Use the results to decide which model or prompt is best for your use case.

### How to add Audio Input in Bulk Evaluations?

{% embed url="<https://youtu.be/w1mKxxIWrRc>" %}
