Data Sources Configuration
Configure how AI/Run CodeMie processes and indexes different types of data sources. This configuration controls chunking strategies, batch processing, and file handling for optimal AI assistant performance.
Overview
Data source loaders control how content from various sources is processed and made available to AI assistants. Each loader is optimized for specific content types and can be tuned for your organization's needs.
The default configuration works for most deployments. Customize these settings if you need to:
- Adjust performance for large-scale data processing
- Fine-tune chunking for specific content types
- Add support for custom file extensions
- Optimize token usage for your LLM models
Configuration Steps
1. Edit Values File
Open codemie-helm-charts/codemie-api/values.yaml and add the configuration blocks below.
2. Add ConfigMap
Add the data sources configuration as a ConfigMap in the extraObjects section:
extraObjects:
- apiVersion: v1
kind: ConfigMap
metadata:
name: datasources-config
data:
datasources-config.yaml: |
---
loaders:
# Loader configurations (see below)
storage:
# Storage configurations (see below)
3. Mount Configuration
Add volume and volume mount configurations:
extraVolumes: |
- name: datasources-config
configMap:
name: datasources-config
extraVolumeMounts: |
- name: datasources-config
mountPath: /app/config/datasources/datasources-config.yaml
subPath: datasources-config.yaml
4. Apply Changes
Deploy the updated configuration:
helm upgrade --install codemie-api \
oci://europe-west3-docker.pkg.dev/or2-msq-epmd-edp-anthos-t1iylu/helm-charts/codemie \
--version x.y.z \
--namespace "codemie" \
-f "./codemie-api/values.yaml" \
--wait --timeout 600s
Replace x.y.z with your version.
Loader Configurations
Code Loader
Processes code files from Git repositories and other code sources. Supports language-aware splitting for better context preservation.
code_loader:
languages_for_splitting:
cpp:
- .cpp
- .h
- .hpp
- .cxx
- .cc
- .C
- .c++
go:
- .go
java:
- .java
js:
- .js
php:
- .php
- .phtml
- .php3
- .php4
- .php5
- .php7
- .phps
- .phpt
proto:
- .proto
python:
- .py
- .pyc
- .pyd
- .pyo
- .pyw
- .pyz
rst:
- .rst
ruby:
- .rb
- .rbx
- .rjs
- .rhtml
- .ru
rust:
- .rs
scala:
- .scala
swift:
- .swift
markdown:
- .md
- .markdown
latex:
- .tex
html:
- .html
- .htm
- .shtml
- .xhtml
sol:
- .sol
chunk_size: 2000 # Characters per chunk
tokens_size_limit: 2000 # Maximum tokens per chunk
chunk_overlap: 30 # Overlap between chunks (characters)
summarization_max_tokens_limit: 4000 # Token limit for summarization
summarization_tokens_overlap: 100 # Overlap for summarization chunks
summarization_batch_size: 10 # Files processed per batch
loader_batch_size: 250 # Documents per processing batch
enable_multiprocessing: false # Enable parallel processing
excluded_extensions:
common:
- .ico
- .mng
- .bpm
- .exe
- .dll
- .jar
- .key
- .mp3
- .mp4
- .otf
- .pyc
- .rar
- .rtf
- .tar
- .gz
- .webm
- .zip
- .xls
- .lock
docs_only:
- .md
- .toml
- .json
- .pdf
- .xlsx
code_only: []
Key Parameters:
chunk_size- Larger chunks provide more context but use more tokenschunk_overlap- Prevents context loss at chunk boundariesloader_batch_size- Higher values improve throughput but use more memoryexcluded_extensions- Skip binary and non-text files
Jira Loader
Processes Jira issues and associated content.
jira_loader:
chunk_size: 1000 # Characters per chunk
chunk_overlap: 50 # Overlap between chunks
loader_batch_size: 50 # Issues per batch
JSON Loader
Processes structured JSON data.
json_loader:
chunk_size: 2000 # Characters per chunk
chunk_overlap: 100 # Overlap between chunks
Confluence Loader
Processes Confluence pages and spaces.
confluence_loader:
loader_max_pages: 1000 # Number of pages lazy_load holds in memory at once before yielding and moving to the next chunk
loader_pages_per_request: 20 # Number of pages returned by a single Confluence API HTTP request (the ?limit= param)
loader_batch_size: 50 # Number of Documents passed to one _process_batch call (splitting + embedding + ES write)
loader_timeout: 180 # Request timeout (seconds)
Key Parameters:
loader_max_pages- Controls the in-memory page buffer for lazy loading; set lower for memory-constrained environmentsloader_pages_per_request- Maps directly to the?limit=parameter of the Confluence API; lower values reduce individual request sizeloader_batch_size- Number of documents sent through splitting, embedding, and Elasticsearch write in a single batchloader_timeout- Increase for slow networks or large pages
File Loader
Processes uploaded files and documents.
file_loader:
chunk_size: 1500 # Characters per chunk
chunk_overlap: 100 # Overlap between chunks
Azure DevOps Wiki Loader
Processes Azure DevOps wiki pages.
azure_devops_wiki_loader:
chunk_size: 1000 # Characters per chunk
chunk_overlap: 50 # Overlap between chunks
loader_batch_size: 50 # Documents per processing batch
Azure DevOps Work Item Loader
Processes Azure DevOps work items, optionally including comments and attachments.
azure_devops_work_item_loader:
chunk_size: 1000 # Characters per chunk
chunk_overlap: 50 # Overlap between chunks
loader_batch_size: 50 # Work items per processing batch
index_comments: true # Index work item comments
index_attachments: true # Index work item attachments
Key Parameters:
index_comments- Enable to include work item comment threads in the indexindex_attachments- Enable to include attached files in the index
Xray Loader
Processes Xray test management data from Jira.
xray_loader:
chunk_size: 1000 # Characters per chunk
chunk_overlap: 50 # Overlap between chunks
loader_batch_size: 50 # Documents per processing batch
SharePoint Loader
Processes SharePoint documents and libraries via the Microsoft Graph API.
sharepoint_loader:
loader_batch_size: 20 # Documents per processing batch
loader_timeout: 300 # Request timeout (seconds)
chunk_size: 2000 # Characters per chunk
chunk_overlap: 200 # Overlap between chunks
max_file_size_mb: 50 # Maximum file size to process (MB)
max_retries: 3 # Retry attempts for failed requests
graph_api_version: "v1.0" # Microsoft Graph API version
graph_base_url: "https://graph.microsoft.com" # Microsoft Graph API base URL
Key Parameters:
max_file_size_mb- Files exceeding this limit are skipped during indexinggraph_api_version- Update if a newer stable Graph API version is requiredmax_retries- Increase for unstable network connections to SharePoint
SVN Loader
Processes Subversion repositories.
svn_loader:
loader_batch_size: 250 # Documents per processing batch
checkout_timeout_seconds: 300 # Timeout for SVN checkout operations
max_file_size_kb: 5000 # Maximum file size to process (KB)
Key Parameters:
checkout_timeout_seconds- Increase for large repositories or slow SVN serversmax_file_size_kb- Files exceeding this limit are skipped during indexing
Storage Configuration
Configure how processed data is stored and indexed in Elasticsearch.
storage:
embeddings_max_docs_count: 20 # Max documents for embedding context
indexing_bulk_max_chunk_bytes: 104857600 # Max bulk request size (100 MB)
indexing_max_retries: 2 # Retry attempts for failed indexing
indexing_error_retry_wait_min_seconds: 10 # Minimum retry wait time (seconds)
indexing_error_retry_wait_max_seconds: 120 # Maximum retry wait time (seconds)
indexing_threads_count: 20 # Parallel indexing threads
processed_documents_threshold: 1000 # Max processed documents stored in Elasticsearch
stale_indexing_threshold_seconds: 300 # Time after which an indexing task is considered stale
stale_indexing_resume_batch_size: 5 # Number of stale tasks resumed per cycle
indexing_heartbeat_interval: 10 # Frequency (in completed docs) for committing indexing stats
Key Parameters:
indexing_threads_count- Increase for faster indexing on high-performance clustersindexing_bulk_max_chunk_bytes- Adjust based on Elasticsearch cluster capacityindexing_max_retries- Worst-case retry duration equalsindexing_max_retries × indexing_error_retry_wait_max_seconds; keep belowstale_indexing_threshold_secondsstale_indexing_threshold_seconds- Tasks that exceed this duration without a heartbeat are treated as stale and resumedindexing_heartbeat_interval- Lower values keepupdate_datefresher and reduce false stale detection
Complete Configuration Example
Full datasources-config.yaml example
extraObjects:
- apiVersion: v1
kind: ConfigMap
metadata:
name: datasources-config
data:
datasources-config.yaml: |
---
loaders:
code_loader:
languages_for_splitting:
cpp:
- .cpp
- .h
- .hpp
- .cxx
- .cc
- .C
- .c++
go:
- .go
java:
- .java
js:
- .js
php:
- .php
- .phtml
- .php3
- .php4
- .php5
- .php7
- .phps
- .phpt
proto:
- .proto
python:
- .py
- .pyc
- .pyd
- .pyo
- .pyw
- .pyz
rst:
- .rst
ruby:
- .rb
- .rbx
- .rjs
- .rhtml
- .ru
rust:
- .rs
scala:
- .scala
swift:
- .swift
markdown:
- .md
- .markdown
latex:
- .tex
html:
- .html
- .htm
- .shtml
- .xhtml
sol:
- .sol
chunk_size: 2000
tokens_size_limit: 2000
chunk_overlap: 30
summarization_max_tokens_limit: 4000
summarization_tokens_overlap: 100
summarization_batch_size: 10
loader_batch_size: 250
enable_multiprocessing: false
excluded_extensions:
common:
- .ico
- .mng
- .bpm
- .exe
- .dll
- .jar
- .key
- .mp3
- .mp4
- .otf
- .pyc
- .rar
- .rtf
- .tar
- .gz
- .webm
- .zip
- .xls
- .lock
docs_only:
- .md
- .toml
- .json
- .pdf
- .xlsx
code_only: []
jira_loader:
chunk_size: 1000
chunk_overlap: 50
loader_batch_size: 50
json_loader:
chunk_size: 2000
chunk_overlap: 100
confluence_loader:
loader_max_pages: 1000
loader_pages_per_request: 20
loader_batch_size: 50
loader_timeout: 180
file_loader:
chunk_size: 1500
chunk_overlap: 100
azure_devops_wiki_loader:
chunk_size: 1000
chunk_overlap: 50
loader_batch_size: 50
azure_devops_work_item_loader:
chunk_size: 1000
chunk_overlap: 50
loader_batch_size: 50
index_comments: true
index_attachments: true
xray_loader:
chunk_size: 1000
chunk_overlap: 50
loader_batch_size: 50
sharepoint_loader:
loader_batch_size: 20
loader_timeout: 300
chunk_size: 2000
chunk_overlap: 200
max_file_size_mb: 50
max_retries: 3
graph_api_version: "v1.0"
graph_base_url: "https://graph.microsoft.com"
svn_loader:
loader_batch_size: 250
checkout_timeout_seconds: 300
max_file_size_kb: 5000
storage:
embeddings_max_docs_count: 20
indexing_bulk_max_chunk_bytes: 104857600
indexing_max_retries: 2
indexing_error_retry_wait_min_seconds: 10
indexing_error_retry_wait_max_seconds: 120
indexing_threads_count: 20
processed_documents_threshold: 1000
stale_indexing_threshold_seconds: 300
stale_indexing_resume_batch_size: 5
indexing_heartbeat_interval: 10