Anaconda Certified Professional (ACP) — Questions and Answers
Question 1: What file is required at the root of a conda package recipe to define its build metadata and dependencies?
- build.sh
- meta.yaml (Correct answer)
- setup.py
- conda.yaml
Correct answer: meta.yaml
The meta.yaml file is the recipe descriptor that defines package name, version, source, build requirements, and runtime dependencies for conda-build.
Question 2: In an ETL pipeline using pandas, which operation is used to combine two DataFrames based on a shared key column (similar to a SQL JOIN)?
- pd.merge() (Correct answer)
- pd.join_tables()
- pd.concat()
- pd.combine()
Correct answer: pd.merge()
`pd.merge()` performs database-style joins between DataFrames on one or more key columns, supporting inner, outer, left, and right join types.
Question 3: How does a histogram differ from a bar chart?
- Bar charts are only for numerical data
- Histograms show distributions; bar charts show categories (Correct answer)
- They are identical
- Histograms are only used in AI
Correct answer: Histograms show distributions; bar charts show categories
A histogram is used for showing the distribution of continuous data, whereas a bar chart represents categorical data.
Question 4: After building a conda package locally, which flag do you use with `conda install` to install it directly from the local build output directory?
- --use-local (Correct answer)
- --file
- --offline
- --local
Correct answer: --use-local
The `--use-local` flag instructs conda to search the local package cache (typically `~/anaconda3/conda-bld/`) before checking remote channels.
Question 5: Which method in pandas handles missing values by filling them with a specified value or strategy (e.g., forward fill)?
- df.dropna()
- df.fillna() (Correct answer)
- df.impute()
- df.replace_nan()
Correct answer: df.fillna()
`df.fillna(value)` replaces NaN values with a constant or uses methods like `ffill` (forward fill) and `bfill` (backward fill) to propagate adjacent valid values.
Question 6: When running `conda audit` on an environment, what output indicates that a package has a critical severity vulnerability?
- A CRITICAL severity tag with the associated CVE ID (Correct answer)
- A broken-pipe error in the terminal
- A yellow WARNING label
- An asterisk (*) next to the package name
Correct answer: A CRITICAL severity tag with the associated CVE ID
`conda audit` outputs vulnerability entries with severity labels (LOW, MEDIUM, HIGH, CRITICAL) alongside the CVE identifier for traceability.
Question 7: What is the purpose of the `.condarc` `allowlist_channels` (formerly `whitelist_channels`) key?
- Caches the listed channels for offline use
- Lists channels that are always searched last
- Marks channels as read-only
- Restricts conda to only use the listed channels, blocking any others (Correct answer)
Correct answer: Restricts conda to only use the listed channels, blocking any others
`allowlist_channels` enforces that conda will only communicate with the explicitly listed channels, preventing use of unauthorized or potentially malicious package sources.
Question 8: Which Python library provides the `Parquet` file format support for high-performance columnar storage of DataFrames?
- feather
- h5py
- openpyxl
- pyarrow (Correct answer)
Correct answer: pyarrow
`pyarrow` (and `fastparquet`) enables pandas to read/write Parquet files via `df.to_parquet()` and `pd.read_parquet()`, offering efficient columnar compression for large datasets.
Question 9: Which pandas method efficiently removes duplicate rows from a DataFrame, keeping only the first occurrence by default?
- df.remove_duplicates()
- df.drop_duplicates() (Correct answer)
- df.deduplicate()
- df.unique_rows()
Correct answer: df.drop_duplicates()
`df.drop_duplicates()` returns a DataFrame with duplicate rows removed, with `keep='first'` as default and options for `keep='last'` or `keep=False` to drop all duplicates.
Question 10: Which technique is commonly used to train AI models?
- Supervised learning (Correct answer)
- Random sampling
- Automated data sorting
- Hard coding patterns
Correct answer: Supervised learning
Supervised learning is a machine learning technique where models are trained using labeled datasets to make predictions.
Question 11: What is the recommended frequency for reviewing and updating package management & environment configuration protocols?
- Relying on periodic external audits as the sole evaluation method
- Tracking activity volume without measuring quality
- Reviewing results only at year-end
- Monitoring outcomes through regular data collection and trend analysis (Correct answer)
Correct answer: Monitoring outcomes through regular data collection and trend analysis
Monitoring outcomes through regular data collection and trend analysis is the correct approach because effective package management & environment configuration in the anaconda certified professional field requires adherence to professional standards, evidence-based practices, and systematic methodology. This approach ensures consistent, high-quality outcomes while maintaining professional accountability.
Question 12: What is the most common mistake professionals make when implementing machine learning & ai integration strategies?
- Creating contingency plans for every possible scenario regardless of probability
- Developing contingency plans for high-probability risk scenarios (Correct answer)
- Responding to problems only after they occur
- Transferring all risk to external partners through contracts
Correct answer: Developing contingency plans for high-probability risk scenarios
Developing contingency plans for high-probability risk scenarios is the correct approach because effective machine learning & ai integration in the anaconda certified professional field requires adherence to professional standards, evidence-based practices, and systematic methodology. This approach ensures consistent, high-quality outcomes while maintaining professional accountability.
Question 13: In pandas, which method converts a column's data type (e.g., from object/string to numeric or datetime)?
- df['col'].astype() (Correct answer)
- df['col'].cast()
- df['col'].to_type()
- df['col'].convert()
Correct answer: df['col'].astype()
`df['col'].astype(dtype)` casts a Series to the specified dtype (e.g., `int64`, `float32`, `str`), and `pd.to_datetime()` / `pd.to_numeric()` handle specialized conversions.
Question 14: Which Python library is commonly used for data visualization?
- TensorFlow
- Matplotlib (Correct answer)
- Flask
- Scikit-learn
Correct answer: Matplotlib
Matplotlib is a widely used Python library that provides tools for creating static, animated, and interactive visualizations.
Question 15: When processing a large dataset in chunks to avoid memory overflow, which pandas parameter in `read_csv()` controls how many rows are loaded per iteration?
- max_rows
- chunksize (Correct answer)
- batch_size
- nrows
Correct answer: chunksize
Setting `chunksize=N` in `pd.read_csv()` returns a `TextFileReader` iterator where each iteration yields a DataFrame of N rows.
Question 16: In the context of anaconda certified professional, which principle most directly governs machine learning & ai integration practices?
- Following popular trends without evaluating their applicability
- Using trial-and-error without systematic documentation
- Relying exclusively on vendor-provided solutions
- Applying evidence-based methodologies with peer-reviewed support (Correct answer)
Correct answer: Applying evidence-based methodologies with peer-reviewed support
Applying evidence-based methodologies with peer-reviewed support is the correct approach because effective machine learning & ai integration in the anaconda certified professional field requires adherence to professional standards, evidence-based practices, and systematic methodology. This approach ensures consistent, high-quality outcomes while maintaining professional accountability.
Question 17: Which command is used to scan a conda environment for known security vulnerabilities?
- conda audit (Correct answer)
- conda check --vulnerabilities
- conda scan --security
- conda verify --cve
Correct answer: conda audit
`conda audit` is the built-in Anaconda tool that scans packages in an environment against known CVE databases.
Question 18: Which conda-build command is used to build a package from a local recipe directory called `my-package`?
- conda create my-package
- conda install my-package
- conda-build my-package (Correct answer)
- conda package my-package
Correct answer: conda-build my-package
Running `conda-build my-package` processes the recipe in that directory and produces a .tar.bz2 or .conda artifact.
Question 19: Which Python library is commonly used for data analysis?
- Seaborn
- Pandas (Correct answer)
- TensorFlow
- Matplotlib
Correct answer: Pandas
Pandas is a powerful Python library used for data manipulation and analysis, providing data structures like DataFrames and Series.
Question 20: In Anaconda Enterprise, which feature restricts users to installing packages only from approved internal channels?
- Package allowlisting via channel configuration (Correct answer)
- conda lock enforcement
- Channel pinning
- Namespace isolation
Correct answer: Package allowlisting via channel configuration
Package allowlisting via channel configuration ensures only vetted, approved packages from internal mirrors can be installed, enforcing compliance.
Question 21: What does the `conda convert` command allow you to do with an existing conda package?
- Convert a .conda file to .whl format
- Repackage a built artifact for a different platform/OS (Correct answer)
- Merge two packages into one
- Convert pip packages to conda format
Correct answer: Repackage a built artifact for a different platform/OS
`conda convert` repackages a compiled conda artifact for alternative platforms (e.g., from linux-64 to osx-64) when no compiled C extensions are involved.
Question 22: What is a neural network in AI?
- A set of static rules for automation
- A manual programming approach
- A predefined list of conditions
- A network of connected nodes for deep learning (Correct answer)
Correct answer: A network of connected nodes for deep learning
A neural network is a computational model inspired by the human brain that is used in deep learning for tasks like image recognition and natural language processing.
Question 23: What is the most common mistake professionals make when implementing package management & environment configuration strategies?
- Creating contingency plans for every possible scenario regardless of probability
- Developing contingency plans for high-probability risk scenarios (Correct answer)
- Responding to problems only after they occur
- Transferring all risk to external partners through contracts
Correct answer: Developing contingency plans for high-probability risk scenarios
Developing contingency plans for high-probability risk scenarios is the correct approach because effective package management & environment configuration in the anaconda certified professional field requires adherence to professional standards, evidence-based practices, and systematic methodology. This approach ensures consistent, high-quality outcomes while maintaining professional accountability.
Question 24: In meta.yaml, which top-level key defines the scripts (build.sh on Linux/macOS and bld.bat on Windows) used to compile and install the package?
- source
- build (Correct answer)
- test
- about
Correct answer: build
The `build` section in meta.yaml controls the build scripts, number (build string), skip conditions, and entry points for the package.
Question 25: Which tool or methodology is most appropriate for analyzing package management & environment configuration outcomes?
- Adjusting boundaries based on individual situations without guidelines
- Maintaining professional boundaries while building collaborative relationships (Correct answer)
- Prioritizing relationships over professional standards
- Maintaining strict formality that inhibits collaboration
Correct answer: Maintaining professional boundaries while building collaborative relationships
Maintaining professional boundaries while building collaborative relationships is the correct approach because effective package management & environment configuration in the anaconda certified professional field requires adherence to professional standards, evidence-based practices, and systematic methodology. This approach ensures consistent, high-quality outcomes while maintaining professional accountability.
Question 26: A new regulation impacts python programming & data science procedures. What should a ACP professional do first?
- Delegating compliance oversight to administrative staff
- Ensuring compliance with current regulatory requirements and standards (Correct answer)
- Complying only with regulations that have enforcement mechanisms
- Interpreting regulations loosely to allow maximum flexibility
Correct answer: Ensuring compliance with current regulatory requirements and standards
Ensuring compliance with current regulatory requirements and standards is the correct approach because effective python programming & data science in the anaconda certified professional field requires adherence to professional standards, evidence-based practices, and systematic methodology. This approach ensures consistent, high-quality outcomes while maintaining professional accountability.
Question 27: In pandas, which method is used to apply a custom function to every row or column of a DataFrame?
- df.map()
- df.apply() (Correct answer)
- df.transform()
- df.execute()
Correct answer: df.apply()
`df.apply()` applies a function along an axis (rows with `axis=1`, columns with `axis=0`), enabling custom transformations across the entire DataFrame.
Question 28: Which conda configuration setting allows an organization to mirror Anaconda's default channel on a private server for security compliance?
- offline_mode
- default_channels (Correct answer)
- proxy_servers
- channel_mirror
Correct answer: default_channels
The `default_channels` setting can be overridden to point to an internal mirror, so all package requests go through a controlled, audited repository.
Question 29: Which conda command exports a complete list of all packages in the current environment to a file for reproducibility?
- conda save environment.yml
- conda freeze > requirements.txt
- conda list --export > packages.txt
- conda env export > environment.yml (Correct answer)
Correct answer: conda env export > environment.yml
`conda env export > environment.yml` captures all packages, versions, and channels in the active environment into a YAML file that can recreate the environment elsewhere.
Question 30: Which command checks a built conda package for common packaging issues such as missing files or incorrect metadata?
- conda audit
- conda-build --check
- conda inspect (Correct answer)
- conda verify
Correct answer: conda inspect
`conda inspect` provides subcommands like `conda inspect linkages` and `conda inspect objects` to audit installed packages for linking issues and metadata correctness.
Question 31: When uploading a package to Anaconda.org via the CLI, which command is used?
- pip upload
- anaconda upload (Correct answer)
- conda upload
- conda push
Correct answer: anaconda upload
The `anaconda upload` command (from the anaconda-client package) pushes a built .tar.bz2 or .conda file to your Anaconda.org channel.
Question 32: What is the recommended frequency for reviewing and updating data visualization & analysis protocols?
- Reviewing results only at year-end
- Monitoring outcomes through regular data collection and trend analysis (Correct answer)
- Relying on periodic external audits as the sole evaluation method
- Tracking activity volume without measuring quality
Correct answer: Monitoring outcomes through regular data collection and trend analysis
Monitoring outcomes through regular data collection and trend analysis is the correct approach because effective data visualization & analysis in the anaconda certified professional field requires adherence to professional standards, evidence-based practices, and systematic methodology. This approach ensures consistent, high-quality outcomes while maintaining professional accountability.
Question 33: When building a data pipeline, what does the term 'idempotency' mean in the context of pipeline task execution?
- A task can process data from any source format
- A task automatically retries on failure
- Running a task multiple times produces the same result as running it once (Correct answer)
- A task runs in parallel with zero overhead
Correct answer: Running a task multiple times produces the same result as running it once
An idempotent pipeline task can be safely re-executed without side effects — running it once or ten times yields the same final state, which is critical for reliable data engineering.
Question 34: In a conda recipe, where would you add a `run_test.py` script to verify the package works after installation?
- In the test section of meta.yaml or as a run_test.py file in the recipe directory (Correct answer)
- In a separate test-requirements.txt file
- In the build section of meta.yaml
- In the source section of meta.yaml
Correct answer: In the test section of meta.yaml or as a run_test.py file in the recipe directory
The `test` section in meta.yaml (or a `run_test.py` file alongside the recipe) specifies imports, commands, and scripts that conda-build runs to validate the installed package.
Question 35: What is the recommended practice to prevent supply chain attacks when using conda?
- Disable SSL verification to speed up downloads
- Always use the `--force-reinstall` flag
- Pin package versions and restrict channels to trusted internal mirrors (Correct answer)
- Use `conda update --all` before every project run
Correct answer: Pin package versions and restrict channels to trusted internal mirrors
Pinning exact package versions and sourcing only from vetted internal mirrors reduces the risk of a malicious package being silently introduced into the environment.
Question 36: Which pandas method reads a CSV file into a DataFrame and can handle large files by specifying a `chunksize` parameter?
- pd.read_table()
- pd.load_csv()
- pd.read_csv() (Correct answer)
- pd.from_csv()
Correct answer: pd.read_csv()
`pd.read_csv()` is the standard function for loading CSV data into a DataFrame, and its `chunksize` parameter returns an iterator of DataFrame chunks for memory-efficient processing.
Question 37: Which pandas method stacks a DataFrame from wide format (one column per variable) to long format (one row per observation)?
- df.stack()
- df.pivot()
- df.melt() (Correct answer)
- pd.wide_to_long()
Correct answer: df.melt()
`df.melt()` unpivots a DataFrame from wide to long format by converting specified columns into rows, creating `variable` and `value` columns.
Question 38: In a meta.yaml recipe, where do you specify packages that are needed ONLY during the build process (e.g., compilers)?
- requirements: host
- requirements: test
- requirements: run
- requirements: build (Correct answer)
Correct answer: requirements: build
The `requirements: build` section lists cross-compilation tools and compilers that run on the build machine but are not needed in the final environment.
Question 39: Which command lists all installed packages in a Conda environment?
- pip freeze
- conda list (Correct answer)
- conda show
- conda packages
Correct answer: conda list
The `conda list` command displays all installed packages within the active Conda environment.
Question 40: Which command is used to create a new virtual environment in Anaconda?
- virtualenv new
- pip install
- conda create (Correct answer)
- python -m venv
Correct answer: conda create
The `conda create` command allows users to create isolated virtual environments with specific dependencies.
Question 41: What is the purpose of the `build_number` field in a meta.yaml recipe?
- Distinguishes multiple builds of the same package version (Correct answer)
- Sets the Python version for the build
- Specifies the conda-build version required
- Defines the number of parallel build jobs
Correct answer: Distinguishes multiple builds of the same package version
The `build_number` increments when a recipe is rebuilt without changing the package version, allowing conda to distinguish between different builds of the same release.
Question 42: What is the purpose of a Conda environment?
- To manage dependencies and isolate projects (Correct answer)
- To uninstall packages globally
- To delete Python installations
- To force package updates system-wide
Correct answer: To manage dependencies and isolate projects
Conda environments help users manage dependencies, prevent conflicts, and isolate project-specific packages.
Question 43: What is a `.conda` file format compared to the older `.tar.bz2` conda package format?
- An encrypted package for enterprise distribution
- A format that only stores pure-Python packages
- A format exclusive to Windows platforms
- A zip-based format with separate metadata and data archives for faster extraction (Correct answer)
Correct answer: A zip-based format with separate metadata and data archives for faster extraction
The `.conda` format is a zip archive containing separate `pkg-*.tar.zst` (data) and `info-*.tar.zst` (metadata) components, enabling faster installs by extracting only needed parts.
Question 44: In Anaconda's role-based access control (RBAC), which role typically has permission to publish packages to a private channel?
- Viewer
- Anonymous user
- Read-only collaborator
- Contributor or higher (e.g., Owner) (Correct answer)
Correct answer: Contributor or higher (e.g., Owner)
In Anaconda's RBAC model, Contributor or Owner roles have the necessary permissions to upload and publish packages to private organizational channels.
Question 45: Which Python library is specifically designed for defining, scheduling, and monitoring data pipeline workflows as Directed Acyclic Graphs (DAGs)?
- Apache Airflow
- Luigi
- All of the above (Correct answer)
- Prefect
Correct answer: All of the above
Luigi, Apache Airflow, and Prefect are all Python-native workflow orchestration frameworks that model pipelines as DAGs with scheduling and monitoring capabilities.
Question 46: To install a package from a private Anaconda.org channel owned by user `myorg`, which flag do you add to `conda install`?
- --private myorg
- --org myorg
- --repo myorg
- -c myorg (Correct answer)
Correct answer: -c myorg
The `-c myorg` (or `--channel myorg`) flag tells conda to search the `myorg` channel on Anaconda.org before the default channels.
Question 47: A new regulation impacts data visualization & analysis procedures. What should a ACP professional do first?
- Ensuring compliance with current regulatory requirements and standards (Correct answer)
- Interpreting regulations loosely to allow maximum flexibility
- Complying only with regulations that have enforcement mechanisms
- Delegating compliance oversight to administrative staff
Correct answer: Ensuring compliance with current regulatory requirements and standards
Ensuring compliance with current regulatory requirements and standards is the correct approach because effective data visualization & analysis in the anaconda certified professional field requires adherence to professional standards, evidence-based practices, and systematic methodology. This approach ensures consistent, high-quality outcomes while maintaining professional accountability.
Question 48: Which file format does `conda audit` primarily reference to identify vulnerable package versions?
- conda-lock.yml
- environment.yml
- NIST NVD / OSV database feeds (Correct answer)
- requirements.txt
Correct answer: NIST NVD / OSV database feeds
`conda audit` queries vulnerability databases such as the NIST National Vulnerability Database (NVD) and OSV to match installed packages against known CVEs.
Question 49: When a conda package recipe uses `{{ version }}` in meta.yaml, where is the value of `version` typically sourced from?
- From a .version file in the source directory
- Automatically from PyPI
- From a set statement at the top of meta.yaml using Jinja2 templating (Correct answer)
- From an environment variable named VERSION
Correct answer: From a set statement at the top of meta.yaml using Jinja2 templating
Jinja2 `{% set version = '1.2.3' %}` at the top of meta.yaml assigns the variable, which is then referenced as `{{ version }}` throughout the recipe for DRY versioning.
Question 50: When automating a data pipeline on a schedule using cron or a task scheduler, which Python standard library module provides programmatic access to run shell commands and subprocesses?
- shutil
- threading
- subprocess (Correct answer)
- os.system only
Correct answer: subprocess
The `subprocess` module provides `subprocess.run()` and `Popen` for launching external processes, capturing output, and handling errors within automated Python pipeline scripts.
Question 51: Which Python library provides the `Pipeline` class to chain preprocessing steps and a final estimator into a single reusable workflow object?
- pandas
- numpy
- scikit-learn (Correct answer)
- scipy
Correct answer: scikit-learn
scikit-learn's `Pipeline` chains transformers and a final estimator so that `fit` and `predict` calls automatically apply all steps in sequence.
Question 52: What is the role of the `source` section in a conda meta.yaml recipe?
- Configures the build machine's source environment
- Specifies where to fetch the package source code (URL, git repo, or local path) (Correct answer)
- Lists the source channels to search for dependencies
- Defines the Python source files to include
Correct answer: Specifies where to fetch the package source code (URL, git repo, or local path)
The `source` section tells conda-build where to download or copy the package source, supporting URLs with checksums, git repositories, and local directory paths.
Question 53: Which SQLAlchemy function is used alongside pandas `read_sql()` to connect to a database and execute SQL queries into a DataFrame?
- make_connection()
- connect_db()
- create_engine() (Correct answer)
- open_session()
Correct answer: create_engine()
`create_engine('dialect+driver://user:pass@host/db')` from SQLAlchemy creates a connection engine that pandas `read_sql()` uses to execute queries and return results as a DataFrame.
Question 54: In pandas, what does `groupby()` followed by `agg()` allow you to do?
- Sort a DataFrame by multiple columns
- Apply multiple aggregation functions to groups simultaneously (Correct answer)
- Filter rows based on group membership
- Join two DataFrames on a common group key
Correct answer: Apply multiple aggregation functions to groups simultaneously
`df.groupby('col').agg({'col1': 'sum', 'col2': 'mean'})` groups rows and applies different aggregation functions to different columns in one vectorized operation.
Question 55: In a data engineering context, what is the purpose of using Dask instead of pandas for large dataset processing in Python?
- Dask enables parallel and out-of-core computation on datasets larger than RAM (Correct answer)
- Dask is only used for streaming real-time data
- Dask provides faster single-threaded operations than pandas
- Dask replaces pandas with a SQL-only interface
Correct answer: Dask enables parallel and out-of-core computation on datasets larger than RAM
Dask partitions large datasets into chunks and processes them in parallel across cores or a cluster, handling data that exceeds available memory with a pandas-compatible API.
Question 56: In pandas, which method writes a DataFrame to a SQL database table using a SQLAlchemy engine?
- df.write_sql()
- df.to_database()
- df.to_sql() (Correct answer)
- df.export_sql()
Correct answer: df.to_sql()
`df.to_sql('table_name', engine, if_exists='replace')` writes DataFrame contents to a SQL table, with options to append, replace, or fail if the table exists.
Question 57: What command renders the final meta.yaml for a recipe after applying all variant substitutions, without actually building it?
- conda-build --dry-run
- conda-build --render
- conda inspect recipe
- conda render (Correct answer)
Correct answer: conda render
`conda render` processes all Jinja2 templating and variant configs in a recipe and outputs the fully resolved meta.yaml without triggering a build.
Question 58: Which security principle is best enforced by using separate conda environments for each project rather than a single shared environment?
- Defense in depth
- Non-repudiation
- Principle of least privilege / isolation (Correct answer)
- Data integrity
Correct answer: Principle of least privilege / isolation
Separate environments enforce isolation and least privilege: a vulnerability or compromised package in one project environment cannot affect others, limiting blast radius.
Question 59: Which file is used to export a Conda environment configuration?
- requirements.txt
- conda.config
- setup.py
- environment.yml (Correct answer)
Correct answer: environment.yml
The `environment.yml` file allows users to export and share Conda environments with dependencies for consistent setups.
Question 60: What does CVE stand for in the context of Anaconda security scanning?
- Common Vulnerabilities and Exposures (Correct answer)
- Conda Version Error
- Common Vulnerability Exposure
- Certified Vulnerability Entry
Correct answer: Common Vulnerabilities and Exposures
CVE stands for Common Vulnerabilities and Exposures, the industry-standard identifier for publicly known security flaws.
Anaconda Certified Professional (ACP)
The Anaconda Certified Professional (ACP) exam validates expertise in the Anaconda data science platform, covering conda package management, data engineering and workflow automation, and machine learning and AI integration using Python.
Exam Rules
- You can skip questions and return to them later
- Flag questions for review before submitting
- No feedback shown until you submit the entire exam
- Unanswered questions count as wrong — answer everything
- 10 pretest questions are mixed in and don't affect your score
- Timer auto-submits when time runs out
- Your progress is auto-saved every 30 seconds