VM4: AI Assistant API (RAG & Web Search)
The assistant service supports search assistant streams, file uploads, and chat.
Assistant Search
- Route:
POST /api/assistant/search - Payload:
{
"query": "Who won the world cup in 2022?",
"mode": "balanced",
"safesearch": 1
} - Flow:
- Retrieves aggregator search results from SearXNG.
- Aggregates search output content context.
- Begins Hono SSE streaming.
- Triggers LLM chain (Groq Llama-3.1 -> Gemini 1.5 -> local Ollama).
SSE Output Stream Format
The server outputs SSE formatted payloads:
citation: Contains search results citation index.token: Individual tokens.[DONE]: Stream completion indicator.
data: {"type": "citation", "sources": [{"title": "2022 FIFA World Cup", "url": "..."}]}
data: {"type": "token", "content": "Argentina "}
data: {"type": "token", "content": "won..."}
data: [DONE]
PDF Chat (RAG)
- File Upload Route:
POST /api/upload- Ingests file (max 10MB PDF), extracts text, generates embeddings (Gemini text-embedding-004), and updates Qdrant collection named
user_vectors_<userId>.
- Ingests file (max 10MB PDF), extracts text, generates embeddings (Gemini text-embedding-004), and updates Qdrant collection named
- Chat Query Route:
POST /api/assistant/chat- Body:
{ "file_id": "...", "query": "..." } - Fetches top matching segments from Qdrant vector index, injects context, and streams responses.
- Body: