โ† All CMA Flashcard Decks

Data Architecture & Management Flashcards

7 cards from real CMA practice questions. Tap to flip, then mark Knew It or Still Learning โ€” missed cards come back until you master them.

Read the first 7 Data Architecture & Management flashcards as text
  1. Which data integration pattern is most appropriate for synchronizing large volumes of data between enterprise systems on a scheduled basis?

    Answer: Batch ETL (Extract, Transform, Load) processing

    Batch ETL is designed to extract, transform, and load large volumes of data at scheduled intervals, making it well-suited for high-volume, time-insensitive data synchronization workloads.

  2. What is the key architectural difference between a data lake and a data warehouse?

    Answer: A data lake stores raw data in native format while a data warehouse stores processed, structured data

    A data lake stores raw data in native format (structured, semi-structured, unstructured) using schema-on-read, while a data warehouse uses schema-on-write with processed, integrated, structured data.

  3. What does the CAP theorem state about distributed data systems?

    Answer: A distributed system cannot simultaneously guarantee Consistency, Availability, and Partition tolerance

    The CAP theorem states that a distributed data system can only guarantee two of three properties simultaneously: Consistency, Availability, and Partition tolerance.

  4. What is the primary advantage of a data mesh architecture approach?

    Answer: Enabling domain teams to own, manage, and serve their data as products

    Data mesh decentralizes data ownership to domain teams who treat data as products, improving scalability, agility, and the application of domain expertise to data quality and management.

  5. Which integration approach best supports real-time data flow between microservices with loose coupling?

    Answer: Event streaming using a message broker such as Apache Kafka

    Event streaming with message brokers like Apache Kafka enables real-time, decoupled data integration through publish-subscribe patterns, ensuring services remain independent and scalable.

  6. What is the primary purpose of a data catalog in enterprise architecture?

    Answer: To provide a searchable inventory of data assets with metadata, lineage, and business context

    A data catalog provides a centralized, searchable inventory of data assets enriched with metadata, lineage, and business context, enabling data discovery and supporting governance.

  7. In data architecture, what is 'data virtualization'?

    Answer: Providing unified, real-time data access across multiple sources without physical data movement

    Data virtualization provides an abstraction layer enabling unified, real-time data access across heterogeneous sources without physically moving or replicating the underlying data.