Files
ai_agent/docs/file-processing.md
furyhawk 8351e73d39 feat: add Zustand stores for conversation, file preview, sidebar, theme, and knowledge base selection
- Implemented `conversation-store` for managing conversations and messages.
- Created `file-preview-store` to handle file preview state.
- Added `sidebar-store` for sidebar visibility management.
- Developed `theme-store` for theme persistence and management.
- Introduced `kb-selection-store` for managing active knowledge base selections with persistence.

chore: define API and chat types

- Added types for API responses, authentication, chat messages, conversations, and projects.
- Defined interfaces for various entities including users, sessions, and message ratings.

build: configure TypeScript and testing setup

- Set up `tsconfig.json` for TypeScript configuration.
- Created `vitest.config.ts` for testing configuration with Vitest.
- Added `vitest.setup.ts` for global test setup including mocks for Next.js router and media queries.
- Configured Vercel deployment settings in `vercel.json`.
2026-06-11 16:54:43 +08:00

7.0 KiB

File Processing

This document covers how files are handled in two contexts: chat file uploads (user-facing) and RAG document ingestion (admin/CLI).

Chat File Uploads

When a user uploads a file in the chat interface, the following pipeline runs:

Flow

1. Upload     POST /api/v1/files/upload
               |
2. Validate    Check MIME type against allowed list + enforce size limit
               |
3. Classify    Determine file_type: "image", "pdf", "docx", "text"
               |
4. Parse       Extract text content (images skip this step)
               |
5. Store       Save file to media/{user_id}/ via FileStorageService
               |
6. Record      Create ChatFile row in database
               |
7. Link        When message is sent, ChatFile is attached via message_id FK
               |
8. Display     Frontend shows images as thumbnails, documents as badges

Supported File Types

Category MIME Types Extensions Processing
Images image/jpeg, image/png, image/webp, image/gif .jpg, .png, .webp, .gif Stored as-is. Sent to LLM as BinaryContent for vision analysis.
PDF application/pdf .pdf Text extracted via configured PDF parser. Appended to prompt as context.
DOCX application/vnd.openxmlformats-officedocument.wordprocessingml.document .docx Paragraphs extracted via python-docx. Appended to prompt as context.
Text text/plain, text/markdown .txt, .md UTF-8 decoded directly. Appended to prompt as context.

PDF Parser Selection (Chat)

PDFs are processed using PyMuPDF. This is a local parser that requires no API key and handles text extraction, table detection, and block-level parsing.

Size Limits

  • Maximum file size: MAX_UPLOAD_SIZE_MB environment variable (default: 50 MB)
  • The limit is enforced server-side after reading the file content.

Storage

Files are saved by FileStorageService to the media/ directory:

media/
  {user_id}/
    document.pdf
    screenshot.png
    ...

ChatFile Model

The ChatFile database model tracks uploaded files:

Field Type Description
id UUID Primary key
user_id UUID/FK Owner (used for access control)
filename String Original filename
mime_type String MIME type (e.g. application/pdf)
size Integer File size in bytes
storage_path String Relative path in storage
file_type String Classified type: image, pdf, docx, text
parsed_content Text Extracted text content (NULL for images)
message_id UUID/FK Linked message (set when message is sent)
created_at DateTime Upload timestamp

Ownership & Access

  • Only the file owner can download their files (GET /files/{id}).
  • The FileUploadService.get_user_file() method compares chat_file.user_id against the requesting user's ID. Returns NotFoundError on mismatch.
  • There is no admin override -- admins cannot access other users' chat files through the file API.

RAG Document Ingestion

When documents are ingested into the RAG knowledge base (via CLI or API), a different pipeline handles parsing, chunking, and embedding.

Ingestion Flow

1. Input       File path (CLI) or uploaded file (API)
                |
2. Parse       DocumentProcessor selects parser by file type
                |
3. Chunk       Text split into segments (configurable size/overlap/strategy)
                |
4. Embed       Chunks embedded via configured provider
                |
5. Store       Vectors written to vector database
                |
6. Track       RAGDocument record created in SQL (status tracking)

Supported Formats

The set of supported formats depends on the configured PDF parser:

Supported file types with the default PyMuPDF parser:

Extension Type Notes
.pdf PDF Text + table extraction via PyMuPDF
.docx Word Paragraph extraction via python-docx
.txt Plain text Direct read
.md Markdown Direct read

PDF Parser Selection (RAG)

RAG ingestion uses PyMuPDF for document parsing (local, no API key required).

Chunking Configuration

Text is split into chunks before embedding. Configure via environment variables:

Variable Default Description
RAG_CHUNK_SIZE 512 Maximum characters per chunk
RAG_CHUNK_OVERLAP 50 Characters of overlap between chunks
RAG_CHUNKING_STRATEGY recursive Strategy: recursive, markdown, fixed

Strategy comparison:

Strategy Best For
recursive General text; splits by paragraph, then sentence, then word
markdown Markdown/structured docs; splits at heading boundaries
fixed Uniform chunk sizes; simplest but may split mid-sentence

Embedding Providers

Embeddings are generated using OpenAI (text-embedding-3-small by default). Set EMBEDDING_MODEL to change the model.

Vector Storage

Vectors are stored in Milvus. Configure with MILVUS_HOST, MILVUS_PORT, MILVUS_DATABASE, and MILVUS_TOKEN.

RAG is Global

Collections are shared across all users:

  • Any authenticated user can search any collection via POST /rag/search or through the AI agent's RAG tool.
  • Only admins can manage collections, upload documents, configure sync sources, and view ingestion logs.
  • There is no per-user document isolation.

Document Tracking

Ingested documents are tracked in the SQL database via the RAGDocument model:

Field Description
collection_name Target collection
filename Original filename
filesize File size in bytes
filetype File extension (without dot)
status processing, done, or error
error_message Error details (if status is error)
vector_document_id ID in the vector store
chunk_count Number of chunks created
storage_path Path to original file (for re-ingestion/download)
created_at Ingestion start time
completed_at Ingestion completion time

Failed ingestions can be retried via POST /rag/documents/{id}/retry.

Sync Operations

Sync operations are tracked via the SyncLog model, recording source, mode, total files, ingested/updated/skipped/failed counts, and timing. View sync history via GET /rag/sync/logs.

Image Description

When processing documents that contain images, the system can optionally describe images using LLM vision capabilities. Set RAG_IMAGE_DESCRIPTION_MODEL to a vision-capable model (defaults to AI_MODEL if empty). The generated descriptions are included in the document text for better semantic search.

Reranking

Search results can optionally be reranked for better relevance. Enable reranking by passing use_reranker=True to the search API. Reranking uses a cross-encoder model (CROSS_ENCODER_MODEL, default: cross-encoder/ms-marco-MiniLM-L6-v2). Runs locally, no API key needed.