RCA Data Loading and Processing 2 — Questions and Answers
Question 1: When configuring a Processing profile in Relativity, what does the 'Deduplication' setting control?
- Whether duplicate documents are merged into a single record based on MD5 hash
- Whether duplicate files are identified and optionally excluded from ingestion based on hash comparison (Correct answer)
- Whether OCR is run on duplicate files to save processing time
- Whether email threading is applied to identify duplicate conversation chains
Correct answer: Whether duplicate files are identified and optionally excluded from ingestion based on hash comparison
Deduplication in a Processing profile identifies files with matching hash values and allows you to exclude duplicates from being ingested as separate documents.
Question 2: Which field in the Processing Source Location settings specifies where Relativity should look for raw data files to process?
- Source Path (Correct answer)
- Network Share Path
- Repository Path
- Collection Path
Correct answer: Source Path
The Source Path in Processing Source Location settings designates the network location where raw evidence files are stored and from which Relativity will pull data for processing.
Question 3: In Relativity Processing, what is the purpose of the 'Inventory' step before full processing?
- To extract text and metadata from all documents immediately
- To generate a preliminary list of files and their metadata without full extraction (Correct answer)
- To create placeholder documents in the workspace for all identified files
- To run virus scans on all source files before ingestion
Correct answer: To generate a preliminary list of files and their metadata without full extraction
Inventory gives you a lightweight scan of the source data to generate file counts and metadata before committing to full processing, allowing better scoping decisions.
Question 4: When using the Relativity Desktop Client (RDC) for a load file import, what does the 'Overlay' mode do?
- Appends new documents to the workspace without touching existing records
- Updates fields on existing documents matched by the overlay identifier without creating duplicates (Correct answer)
- Deletes existing records and replaces them with the incoming data
- Merges metadata from the load file with metadata from processing
Correct answer: Updates fields on existing documents matched by the overlay identifier without creating duplicates
Overlay mode uses a designated identifier field to match incoming rows to existing documents and updates only the specified fields on those matched records.
Question 5: What happens to a Processing job when a file encounters a 'Password Protected' error during text extraction?
- The job fails entirely and must be restarted after the password is supplied
- The file is skipped and flagged with an error; other files continue processing (Correct answer)
- Relativity automatically attempts common passwords before flagging the file
- The file is published to the workspace with a blank extracted text field without flagging
Correct answer: The file is skipped and flagged with an error; other files continue processing
Password-protected files are individually flagged with an error status while the remainder of the processing job continues, allowing you to address them separately.
Question 6: Which Processing profile setting controls whether email metadata fields such as 'To', 'From', and 'Subject' are mapped to workspace fields during publishing?
- Email Header Extraction
- Metadata Field Mapping
- Publishing Profile (Correct answer)
- Field Catalog Mapping
Correct answer: Publishing Profile
The Publishing Profile within Processing settings controls how extracted metadata fields, including email headers, are mapped to destination workspace fields when documents are published.
Question 7: In Relativity, what is the significance of the 'Global Deduplication' option versus 'Custodian Deduplication' during processing?
- Global deduplication removes duplicates across all custodians; custodian deduplication removes duplicates only within each individual custodian's data set (Correct answer)
- Global deduplication runs after publishing; custodian deduplication runs before text extraction
- Global deduplication applies to emails only; custodian deduplication applies to loose files only
- There is no functional difference — both options produce identical results
Correct answer: Global deduplication removes duplicates across all custodians; custodian deduplication removes duplicates only within each individual custodian's data set
Global deduplication treats the entire job as one pool, eliminating cross-custodian duplicates, while custodian deduplication only removes duplicates found within a single custodian's data.
When configuring a Processing profile in Relativity, what does the 'Deduplication' setting control?