Overview
Flow ID: corpus-upload
Category: Dataset/Corpus Management
Estimated Duration: 2-10 minutes (depending on file size)
User Role: All Users
Complexity: Simple
Purpose: This flow allows users to upload a pre-processed dataset file (in JSONL format) to the application. These datasets contain structured information that can be queried during AI conversations, enabling the AI to provide answers based on your specific documents and data rather than just its general knowledge.
Related Flows
- Activate/Deactivate Dataset - Make dataset available for chat queries
- View Dataset Details - Verify uploaded contents
- Chat with Dataset Query Enabled - Use uploaded dataset
- Create New Blockify Job - Alternative: create dataset from documents
- Upload Embedding Model - Add required embedding model
Prerequisites
Before starting, users must have:
- Application installed and running
- Dataset file in JSONL format (each line is a JSON object with "text" and "vector" fields)
- Embedding model available (dataset must have been created with the same embedding model)
- Sufficient disk space for the dataset file
- Knowledge of which embedding model was used to create the dataset
Step-by-Step Flow
Main Path (Happy Path)
Step 1: Navigate to Settings or Datasets
- User Action: Click "Settings" in navigation, OR click "Datasets" if available as separate navigation item
- System Response: Settings page or Datasets page loads
- UI Elements Visible:
- Navigation menu with selected item highlighted
- Page content area
- If Settings: Multiple tabs across top
- If Datasets: Dataset list view
- Visual Cues: Active navigation item highlighted
Step 2: Access Dataset Upload Area
- User Action:
- If in Settings: Look for dataset-related tab or section
- If in Datasets page: Look for "Upload Dataset" or "Add Dataset" button
- System Response: Upload interface becomes available
- UI Elements Visible:
- "Upload Dataset" or "Add Dataset" button (typically prominent, may be blue)
- Possibly: List of existing datasets shown
- Instructions or helper text
- Visual Cues: Upload button stands out with color or icon
Step 3: Initiate Dataset Upload
- User Action: Click "Upload Dataset" or "Add Dataset" button
- System Response: Upload modal or form appears
- UI Elements Visible:
- Modal dialog overlaying current page
- Title: "Upload Dataset" or similar
- File upload area/button
- Text input for "Dataset Name"
- Dropdown or selector for "Embedding Model"
- "Cancel" and "Upload" buttons
- Close button (X) in corner
- Visual Cues:
- Modal centered on screen
- Background slightly dimmed
- Clear form structure
Step 4: Select Dataset File
- User Action: Click "Choose File" or file upload button
- System Response: Operating system file browser opens
- UI Elements Visible:
- Native file browser window
- File navigation interface
- File type filter (may show .jsonl files)
- Visual Cues: Standard OS file browser appearance
Step 5: Navigate to and Select File
- User Action: Browse to dataset file location, select the .jsonl file, click "Open"
- System Response:
- File browser closes
- Modal reappears
- Selected filename displayed
- Dataset name may auto-populate from filename
- UI Elements Visible:
- Selected filename shown in upload area
- Dataset name field populated (or ready for input)
- Embedding model dropdown active
- File size may be displayed
- Visual Cues:
- Filename visible confirms selection
- Form appears complete or nearly complete
Step 6: Enter or Confirm Dataset Name
- User Action: Review auto-populated name or type a custom dataset name
- System Response: Text appears as typed
- UI Elements Visible:
- Text input field with dataset name
- Cursor active in field if editing
- Character count or validation (if applicable)
- Visual Cues:
- Active text field when clicked
- Name should be descriptive
Step 7: Select Embedding Model
- User Action: Click embedding model dropdown and select the model that was used to create this dataset
- System Response: Dropdown expands showing available embedding models
- UI Elements Visible:
- Dropdown list of embedding models
- Model names (e.g., "Jina Embeddings", "BGE-Small")
- Currently selected model highlighted
- Visual Cues:
- Dropdown opens smoothly
- Selected model appears in field when closed
- Note: Critical to select the correct model; mismatched models won't work properly
Step 8: Confirm and Upload
- User Action: Click "Upload" or "Save" button
- System Response:
- Upload begins
- Progress indicator appears
- File transfers to application
- UI Elements Visible:
- Progress bar (animated, filling left to right)
- Percentage indicator (e.g., "Uploading... 35%")
- Status text: "Uploading dataset..."
- Cancel button may be disabled during upload
- Visual Cues:
- Animated progress bar
- Percentage increases
- May show upload speed or time remaining
Step 9: Wait for Upload to Complete
- User Action: Wait while file uploads (time varies by file size)
- System Response:
- Progress continues advancing
- System validates file format
- Dataset is registered in database
- UI Elements Visible:
- Continuing progress updates
- Status may change to "Processing..." after upload completes
- Visual Cues:
- Smooth progress animation
- System is working
Step 10: Upload Completes Successfully
- User Action: No action required
- System Response:
- Progress reaches 100%
- Success message appears
- Modal automatically closes OR shows success state with "Done" button
- Returns to dataset list or settings page
- UI Elements Visible:
- Success message (green, with checkmark): "Dataset uploaded successfully!"
- Dataset now appears in datasets list
- Dataset entry shows: Name, embedding model, size, date
- Visual Cues:
- Green color indicates success
- Checkmark icon
- Smooth transition back to main view
Step 11: Verify Dataset in List
- User Action: Look for newly uploaded dataset in the list
- System Response: Dataset list displays with new entry
- UI Elements Visible:
- Dataset list or table showing all datasets
- New dataset entry with:
- Dataset name
- Embedding model name
- File size
- Upload date/time
- Status indicator (may show "Ready" or "Inactive")
- Action buttons (Activate, View, Download, Delete)
- Visual Cues:
- New dataset may be highlighted or positioned at top
- Clear dataset information displayed
Final Step: Dataset Available for Use
- Success Indicator:
- Dataset appears in datasets list
- No error messages
- Dataset can be activated
- Dataset shows correct embedding model
- System State Change:
- Dataset file stored in application's data directory
- Dataset registered in database
- Dataset available for selection in chat settings
- Dataset can be activated for use in conversations
- Next Possible Actions:
- Activate this dataset for use in chat (see corpus-activation.md)
- View dataset contents to verify data
- Upload additional datasets
- Return to chat and enable dataset query
- Export or manage existing datasets
Alternative Paths & Strategies
Strategy A: Drag and Drop Upload
When to use: If modal supports drag-and-drop file selection
Steps:
- Open upload modal (Steps 1-3)
- Instead of clicking "Choose File", open file browser separately
- Drag dataset file from file browser to upload modal
- Drop file in designated area
- Filename appears, continue from Step 6
Strategy B: Upload from Datasets List Page
When to use: If dedicated Datasets page exists with direct upload
Steps:
- Navigate directly to Datasets page (not through Settings)
- Click "Upload" or "+" button in datasets view
- Upload modal or inline form appears
- Continue with Steps 4-11
Strategy C: Replace Existing Dataset
When to use: User wants to update an existing dataset with new version
Steps:
- Navigate to dataset list
- Find existing dataset to replace
- Click "Delete" or "Remove" on old dataset
- Confirm deletion
- Follow main path to upload new version with same name
Strategy D: Batch Upload Multiple Datasets
When to use: User has several dataset files to upload
Steps:
- Upload first dataset (full main path)
- When modal closes, immediately click "Upload Dataset" again
- Select second dataset file
- Repeat for each dataset
- All datasets appear in list
QA Note: Current interface may require separate uploads. True batch upload not confirmed in knowledge base.
Error States & Recovery
Error 1: Invalid File Format
Cause: File is not in JSONL format or has incorrect structure
User Experience:
- Error message after upload: "Invalid file format" or "File must be JSONL"
- Upload fails
- May indicate specific format issue
Recovery Steps:
- Verify file is .jsonl format
- Open file in text editor to check structure
- Each line should be valid JSON with "text" and "vector" fields
- Vectors should be arrays of numbers
- Fix file format or regenerate from source
- Try uploading again
Error 2: Embedding Model Not Available
Cause: Selected embedding model doesn't exist or isn't loaded
User Experience:
- Error message: "Embedding model not found" or similar
- Cannot complete upload
- Dropdown may show no models
Recovery Steps:
- Cancel upload
- Navigate to Settings > Models
- Upload the embedding model that was used to create this dataset
- Return to dataset upload
- Select newly uploaded embedding model
- Complete upload
Error 3: Embedding Model Mismatch
Cause: Dataset created with different embedding model than selected
User Experience:
- Upload may succeed but dataset won't work properly
- Search results will be poor or nonsensical
- May not show error immediately (discovered during use)
Recovery Steps:
- Delete incorrectly associated dataset
- Determine which embedding model was used to create dataset (check documentation)
- Upload correct embedding model if not available
- Re-upload dataset with correct embedding model selected
QA Note: System may not validate embedding model compatibility at upload time. User discovers issue during use.
Error 4: Insufficient Disk Space
Cause: Not enough free disk space for dataset file
User Experience:
- Error during upload: "Insufficient disk space" or "Upload failed"
- Progress may stop partway
- Upload does not complete
Recovery Steps:
- Free up disk space (delete old files, datasets, or models)
- Check dataset file size before uploading
- Ensure adequate free space (at least 2x file size recommended)
- Try upload again after freeing space
Error 5: Dataset Name Already Exists
Cause: Another dataset with the same name already exists
User Experience:
- Error message: "Dataset name already exists" or "Name must be unique"
- Cannot save
- Upload process stops
Recovery Steps:
- Change dataset name to something unique
- Add version number or date (e.g., "TechDocs-v2" or "TechDocs-2025-01")
- Or delete/rename existing dataset with conflicting name
- Complete upload with unique name
Error 6: File Too Large
Cause: Dataset file exceeds maximum size limit
User Experience:
- Error message: "File too large" or "Exceeds maximum size"
- Upload rejected before starting or fails partway
- May indicate size limit
Recovery Steps:
- Check file size and system limits
- Split dataset into smaller chunks if possible
- Remove less important entries from dataset
- Or upgrade system resources if datasets are legitimately large
QA Note: Practical limit depends on system RAM and disk space, not hard-coded limit.
Error 7: Corrupted File
Cause: File is damaged or incomplete
User Experience:
- Error during processing: "File corrupted" or "Invalid data"
- Upload may complete but validation fails
- Cannot use dataset
Recovery Steps:
- Verify file integrity (check file size matches expected)
- Re-download or re-export file from source
- Try uploading new copy
- If file consistently fails, may need to regenerate from original data
Version History
| Date | Version | Author | Changes |
|---|---|---|---|
| 2025-10-04 | 1.1 | Iternal Technologies | Initial comprehensive documentation |
Notes
Important Considerations:
- Dataset file must be in JSONL format (newline-delimited JSON)
- Each line must have "text" and "vector" fields
- Vector dimensions must match the embedding model used
- Embedding model must be the SAME one used to create the dataset originally
- File is copied to application directory; original can be safely deleted after upload
- Large datasets (>500MB) may take several minutes to upload and process
JSONL Format Example:
Best Practices:
- Use descriptive dataset names that indicate content (e.g., "CompanyPolicy2025", "TechnicalDocs-ProductX")
- Document which embedding model was used when creating datasets
- Keep dataset files backed up externally
- Test small sample datasets before uploading very large ones
- Organize datasets by topic or domain for easier selection
Common User Questions:
- "Where do I get dataset files?" - Create them using Blockify jobs or import from external sources
- "What format should the file be?" - JSONL (JSON Lines) with text and vector fields
- "How do I know which embedding model to use?" - Must match the model used to create the dataset
- "Can I edit a dataset after uploading?" - No, must delete and re-upload; edit the file externally first
- "How large can datasets be?" - Limited by disk space and RAM; very large datasets (GB+) may be slow to search
Trigger
What initiates this flow:
- User manually initiates
Specific trigger: User wants to add a new dataset for AI-powered querying, typically because:
- They have a pre-processed dataset file ready to use
- They want to add knowledge to the AI system
- They've received a dataset from another source
- They want to import previously exported data
- They're testing the dataset query feature
User Intent Analysis
Primary Intent
Import a structured dataset file into the application so it can be used for AI-powered question answering and information retrieval during conversations.
Secondary Intents
- Make specific knowledge available to the AI
- Enable dataset query features
- Build a library of queryable datasets
- Restore previously exported datasets
- Share datasets between instances or users
Subintents
- Ensure dataset is compatible with the application
- Associate dataset with correct embedding model
- Make dataset easily identifiable for future use
- Verify successful upload