Skip to main content

VM4: AI Assistant API (RAG & Web Search)

The assistant service supports search assistant streams, file uploads, and chat.

  • Route: POST /api/assistant/search
  • Payload:
    {
    "query": "Who won the world cup in 2022?",
    "mode": "balanced",
    "safesearch": 1
    }
  • Flow:
    1. Retrieves aggregator search results from SearXNG.
    2. Aggregates search output content context.
    3. Begins Hono SSE streaming.
    4. Triggers LLM chain (Groq Llama-3.1 -> Gemini 1.5 -> local Ollama).

SSE Output Stream Format​

The server outputs SSE formatted payloads:

  • citation: Contains search results citation index.
  • token: Individual tokens.
  • [DONE]: Stream completion indicator.
data: {"type": "citation", "sources": [{"title": "2022 FIFA World Cup", "url": "..."}]}
data: {"type": "token", "content": "Argentina "}
data: {"type": "token", "content": "won..."}
data: [DONE]

PDF Chat (RAG)​

  • File Upload Route: POST /api/upload
    • Ingests file (max 10MB PDF), extracts text, generates embeddings (Gemini text-embedding-004), and updates Qdrant collection named user_vectors_<userId>.
  • Chat Query Route: POST /api/assistant/chat
    • Body: { "file_id": "...", "query": "..." }
    • Fetches top matching segments from Qdrant vector index, injects context, and streams responses.