> For the complete documentation index, see [llms.txt](https://docs.apryse.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.apryse.com/core/augmenting-llms-with-smart-data-extraction/rag-guide.md).

# Using Smart Data Extraction to Augment Contextual LLM Queries - Overview

Learn how to enhance Large-Language Models (LLMs) with structured data using Apryse Data Extraction Module. Set up your environment to run examples efficiently. The Apryse Server SDK streamlines secur

## Overview

Large-Language Models (LLMs) are a powerful tool for a wide range of tasks, including information retrieval, summarization, logical reasoning, and much more. A general limitation of LLMs is that their contextual knowledge is confined to the information available in their training data, meaning that their ability to answer questions about private data or events that occurred after their training will be limited. One workaround for this issue is to provide the contextual data that you are interested in within the query itself, giving the LLM access to the data that was not present in its training set. This guide will show you how to extract structured data from your documents using the [Apryse Data Extraction Module](/core/learn-more/modules.md#data-extraction-module), and will show you several techniques for providing this data as context to an LLM along with a query.

This guide will work with **Open AI's Python API**, although it could be modified to work with other platforms as well.

{% hint style="info" %}
**NOTE**

Some of the operations contained within are long-running or cost money to perform. To save time and money, we use disk caching when appropriate. This allows you to modify the queries without needing to repeat all the other preprocessing. To clear the cache, you will need to delete the associated files.
{% endhint %}

## Get Started

[Setup](/core/augmenting-llms-with-smart-data-extraction/setup.md)

In this section, we show how to set up your environment to run the examples.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.apryse.com/core/augmenting-llms-with-smart-data-extraction/rag-guide.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
