ragflow/docs/guides/chat/best_practices/accelerate_question_answering.mdx

---
sidebar_position: 1
slug: /accelerate_question_answering
---

# Accelerate answering
import APITable from '@site/src/components/APITable';

A checklist to speed up question answering for your chat assistant.

---

Please note that some of your settings may consume a significant amount of time. If you often find that your question answering is time-consuming, here is a checklist to consider:

- Disabling **Multi-turn optimization** will reduce the time required to get an answer from the LLM.
- Leaving the **Rerank model** field empty will significantly decrease retrieval time.
- Disabling the **Reasoning** toggle will reduce the LLM's thinking time. For a model like Qwen3, you also need to add `/no_think` to the system prompt to disable reasoning.
- When using a rerank model, ensure you have a GPU for acceleration; otherwise, the reranking process will be *prohibitively* slow.

:::tip NOTE 
Please note that rerank models are essential in certain scenarios. There is always a trade-off between speed and performance; you must weigh the pros against cons for your specific case.
:::

- Disabling **Keyword analysis** will reduce the time to receive an answer from the LLM.
- When chatting with your chat assistant, click the light bulb icon above the *current* dialogue and scroll down the popup window to view the time taken for each task:  
   ![enlighten](https://github.com/user-attachments/assets/fedfa2ee-21a7-451b-be66-20125619923c)  


```mdx-code-block
<APITable>
```

| Item name         | Description                                                                                   |
| ----------------- | --------------------------------------------------------------------------------------------- |
| Total             | Total time spent on this conversation round, including chunk retrieval and answer generation. |
| Check LLM         | Time to validate the specified LLM.                                                           |
| Create retriever  | Time to create a chunk retriever.                                                             |
| Bind embedding    | Time to initialize an embedding model instance.                                               |
| Bind LLM          | Time to initialize an LLM instance.                                                           |
| Tune question     | Time to optimize the user query using the context of the mult-turn conversation.              |
| Bind reranker     | Time to initialize an reranker model instance for chunk retrieval.                            |
| Generate keywords | Time to extract keywords from the user query.                                                 |
| Retrieval         | Time to retrieve the chunks.                                                                  |
| Generate answer   | Time to generate the answer.                                                                  |

```mdx-code-block
</APITable>
```
Added document: Accelerate document indexing and retrieval (#4600) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-01-24 11:58:15 +08:00			`---`
Docs: Restructured docs (#7614) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-05-13 15:49:08 +08:00			`sidebar_position: 1`
Restructured guides (#5555) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-03-03 17:13:37 +08:00			`slug: /accelerate_question_answering`
Added document: Accelerate document indexing and retrieval (#4600) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-01-24 11:58:15 +08:00			`---`

Added 0.17.0 release notes (#5608) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-03-04 19:21:28 +08:00			`# Accelerate answering`
Added document: Accelerate document indexing and retrieval (#4600) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-01-24 11:58:15 +08:00			`import APITable from '@site/src/components/APITable';`

Docs: How to accelerate question answering (#10179) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-09-19 18:18:46 +08:00			`A checklist to speed up question answering for your chat assistant.`
Added document: Accelerate document indexing and retrieval (#4600) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-01-24 11:58:15 +08:00
			`---`

Restructured guides (#5555) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-03-03 17:13:37 +08:00			`Please note that some of your settings may consume a significant amount of time. If you often find that your question answering is time-consuming, here is a checklist to consider:`
Added document: Accelerate document indexing and retrieval (#4600) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-01-24 11:58:15 +08:00
Docs: How to accelerate question answering (#10179) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-09-19 18:18:46 +08:00			`- Disabling Multi-turn optimization will reduce the time required to get an answer from the LLM.`
			`- Leaving the Rerank model field empty will significantly decrease retrieval time.`
			- Disabling the Reasoning toggle will reduce the LLM's thinking time. For a model like Qwen3, you also need to add `/no_think` to the system prompt to disable reasoning.
UI updates. (#6398) ### What problem does this PR solve? Updated UI descriptions for delimiters and recommended chunk size ### Type of change - [x] Documentation Update 2025-03-21 16:50:20 +08:00			`- When using a rerank model, ensure you have a GPU for acceleration; otherwise, the reranking process will be prohibitively slow.`

			`:::tip NOTE`
			`Please note that rerank models are essential in certain scenarios. There is always a trade-off between speed and performance; you must weigh the pros against cons for your specific case.`
			`:::`

Docs: How to accelerate question answering (#10179) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-09-19 18:18:46 +08:00			`- Disabling Keyword analysis will reduce the time to receive an answer from the LLM.`
Added document: Accelerate document indexing and retrieval (#4600) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-01-24 11:58:15 +08:00			`- When chatting with your chat assistant, click the light bulb icon above the current dialogue and scroll down the popup window to view the time taken for each task:`
			`![enlighten](https://github.com/user-attachments/assets/fedfa2ee-21a7-451b-be66-20125619923c)`


			```mdx-code-block
			`<APITable>`
			```

Added 0.17.0 release notes (#5608) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-03-04 19:21:28 +08:00			`\| Item name \| Description \|`
			`\| ----------------- \| --------------------------------------------------------------------------------------------- \|`
Added document: Accelerate document indexing and retrieval (#4600) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-01-24 11:58:15 +08:00			`\| Total \| Total time spent on this conversation round, including chunk retrieval and answer generation. \|`
Added 0.17.0 release notes (#5608) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-03-04 19:21:28 +08:00			`\| Check LLM \| Time to validate the specified LLM. \|`
			`\| Create retriever \| Time to create a chunk retriever. \|`
			`\| Bind embedding \| Time to initialize an embedding model instance. \|`
			`\| Bind LLM \| Time to initialize an LLM instance. \|`
			`\| Tune question \| Time to optimize the user query using the context of the mult-turn conversation. \|`
			`\| Bind reranker \| Time to initialize an reranker model instance for chunk retrieval. \|`
			`\| Generate keywords \| Time to extract keywords from the user query. \|`
			`\| Retrieval \| Time to retrieve the chunks. \|`
			`\| Generate answer \| Time to generate the answer. \|`
Added document: Accelerate document indexing and retrieval (#4600) ### What problem does this PR solve? ### Type of change - [x] Documentation Update 2025-01-24 11:58:15 +08:00
			```mdx-code-block
			`</APITable>`
			```