AIF-C01 sample questions with answers

10 free practice questions for the AWS Certified AI Practitioner exam. Try each one, then open the answer to see why the right option wins and every other option loses.

Question 1Fundamentals of AI and ML

An online marketplace wants to flag suspicious account registrations and payments. The company employs no data scientists, but its analysts hold a few years of records showing which past registrations turned out to be fraudulent, and they want to build and use a fraud model themselves through a no-code interface on AWS. Which approach should the company use?

  1. A.

    Amazon Personalize, to rank registrations by how closely each one resembles the registrations made by the marketplace's most valuable customers.

  2. B.

    Amazon Comprehend, to analyse the text of each registration form and infer from its wording whether the applicant's intent is fraudulent.

  3. C.

    Amazon SageMaker Canvas, to build a binary classification model from the labelled fraud history without code and generate predictions for new events.

  4. D.

    Amazon Rekognition, to compare the profile photograph supplied at registration against a stored collection of photographs of known fraudsters.

Show answer

Answer: C

Fraud detection from labelled history is a binary classification problem, and SageMaker Canvas lets analysts build and use that model without writing code.

  • A. Personalize generates recommendations from interaction data; it does not detect fraudulent activity.
  • B. Comprehend analyses natural language text and would ignore the behavioural and payment signals.
  • C. Canvas builds a binary classification model from the analysts' own labelled fraud history with no code, which fits both the data and the skills available.
  • D. Face comparison addresses only one narrow signal and presumes a collection of known fraudster photographs.
Question 2Fundamentals of AI and ML

A grocery chain has gathered two years of point-of-sale records for a demand-forecasting project. Before any cleaning or modelling starts, an analyst plots the distribution of every column, counts missing values and reviews correlations between fields. Which stage of the ML pipeline is the analyst performing?

  1. A.

    Model monitoring, to compare live traffic against a training baseline

  2. B.

    Hyperparameter tuning, to search for the training configuration that scores best on validation data

  3. C.

    Model deployment, to make the trained model available behind an endpoint

  4. D.

    Exploratory data analysis, to understand the data before preparing it

Show answer

Answer: D

Profiling distributions, missing values and correlations before cleaning is exploratory data analysis.

  • A. Monitoring compares production traffic with a baseline and requires a deployed endpoint.
  • B. Tuning searches training configurations and can only happen once a model is being trained.
  • C. Deployment publishes a trained model for inference; no model exists at this point.
  • D. Profiling distributions, missing values and correlations is exactly what the EDA stage does.
Question 3Fundamentals of AI and ML

A conservation charity has 3,000 unlabelled camera-trap photographs and wants a model that finds and names six animal species in new photos. Its volunteers can recognise the species but nobody knows ML, and the charity wants one AWS service in which the volunteers draw the bounding boxes and then train the model. Which AWS capability fits this stage?

  1. A.

    Amazon Rekognition Custom Labels, whose console lets volunteers draw labelled bounding boxes and then train a detector

  2. B.

    Amazon Rekognition label detection, which returns general labels such as Animal without any training

  3. C.

    Amazon SageMaker Feature Store, which stores curated features for training and inference

  4. D.

    Amazon Quick Sight, which builds interactive dashboards from business data

Show answer

Answer: A

Rekognition Custom Labels lets the team label images with bounding boxes in its console and then train a custom detector without ML expertise.

  • A. Custom Labels provides console bounding-box labelling and then trains a detector on the labelled images.
  • B. Pre-trained label detection gives general labels and cannot be taught six specific species.
  • C. Feature Store holds engineered feature values, not human annotations of raw images.
  • D. Amazon Quick Sight visualises business data and offers no image annotation workflow.
Question 4Fundamentals of AI and ML

A hospital group's leadership has heard that its data science team plans to 'adopt MLOps' next year, and asks what that will change about how the team's models are built, released and kept running. Which statement BEST describes the goal of MLOps?

  1. A.

    Writing prompts that make a foundation model produce the output format an application expects.

  2. B.

    Choosing the algorithm and hyperparameter values that produce the highest accuracy score on the available training data.

  3. C.

    Collecting and labelling enough training examples to cover every case the model may encounter.

  4. D.

    Applying repeatable, automated and monitored processes so models can be built, deployed and maintained reliably.

Show answer

Answer: D

MLOps applies repeatable, automated and monitored engineering practice to the whole model lifecycle.

  • A. Prompt engineering shapes a foundation model's output and is unrelated to lifecycle operations.
  • B. That describes model training and tuning, one stage inside the lifecycle MLOps governs.
  • C. Data collection and labelling is an early pipeline stage rather than an operational discipline.
  • D. Repeatability, automation, versioning and monitoring across the lifecycle is precisely what MLOps means.
Question 5Fundamentals of AI and ML

A water utility is trying to detect a rare pump-failure signature in raw vibration recordings from its sensors. Over eight months its analysts have hand-engineered 30 features from each recording, such as peak amplitude, the energy in a few frequency bands and simple summary statistics, and have trained classical models on those features. Detection has plateaued at 61% and adding further features of the same kind no longer moves it. An engineer proposes a deep learning model that takes the raw waveform itself as its input. Which statement BEST explains why that proposal might help?

  1. A.

    Deep learning models are not bound by the need for labelled failure examples, so the very small number of recorded failures stops being a constraint on accuracy once the raw waveform is used as the model input.

  2. B.

    The plateau is caused by the 30 features being too many for a classical model to use, and a deep network helps by reducing the number of inputs that it considers down to a much smaller set.

  3. C.

    Deep learning always reaches higher accuracy than classical machine learning on any dataset, so changing the model family is guaranteed to lift detection above the 61% that the current approach reached.

  4. D.

    The signature is a pattern in the raw signal that the 30 hand-designed features do not capture, and a deep network learns its own representation from the waveform instead of depending on features a person thought to define.

Show answer

Answer: D

Deep learning helps here because it learns features from the raw signal that analysts did not design.

  • A. Supervised deep learning needs labelled examples too, and generally more of them than a classical model needs.
  • B. Thirty features is a small input space for any classical model, and deep networks typically expand rather than shrink the representation.
  • C. No model family wins on every dataset; deep learning often loses on small, well-engineered tabular problems.
  • D. Identifies the feature ceiling and the specific property of deep learning, learned representations, that could break through it.
Question 6Fundamentals of AI and ML

A retailer must clean and transform 4 TB of transaction data stored in Amazon S3 before it can be used for model training. The data engineering team wants managed AWS services for this preparation stage rather than servers it maintains itself. Which TWO services are appropriate? (Select TWO.)

Choose 2.

  1. A.

    Amazon CloudWatch, for collecting metrics and logs emitted by the preparation jobs.

  2. B.

    AWS Glue DataBrew, for visual data cleaning and normalisation without writing code.

  3. C.

    Amazon SageMaker Model Registry, for storing approved versions of the trained model.

  4. D.

    Amazon SageMaker Model Monitor, for checking the quality of data arriving at an endpoint.

  5. E.

    AWS Glue, for serverless extract, transform and load jobs over data in Amazon S3.

Show answer

Answer: B, E

AWS Glue and AWS Glue DataBrew are the managed services for preparing large datasets before training.

  • A. CloudWatch observes the jobs but performs no cleaning or transformation of the data itself.
  • B. Glue DataBrew provides visual, code-free cleaning and normalisation of large datasets.
  • C. Model Registry versions trained models and plays no part in the data preparation stage.
  • D. Model Monitor, now closed to new customers, inspects data arriving at a deployed endpoint, not a bulk preparation dataset.
  • E. AWS Glue runs serverless ETL at terabyte scale directly against data stored in Amazon S3.
Question 7Fundamentals of AI and ML

A healthcare analytics company retrains its patient no-show model every month. Today a data scientist runs notebook cells by hand, copies the model file to an S3 bucket and updates the endpoint manually. Two releases were rolled back because nobody could tell which data and code produced the deployed model. Leadership wants a repeatable, auditable process with approval before production and the ability to reproduce any past model. Which approach BEST meets these MLOps goals?

  1. A.

    Schedule the existing notebook with a cron job on an EC2 instance so it runs automatically each month.

  2. B.

    Train once with a larger dataset so the model no longer needs monthly retraining.

  3. C.

    Store the monthly model files in Amazon S3 with versioning enabled and continue updating the endpoint by hand.

  4. D.

    Build an Amazon SageMaker Pipeline that runs preprocessing, training and evaluation steps, registers versioned models in SageMaker Model Registry, and deploys only approved versions.

Show answer

Answer: D

SageMaker Pipelines plus Model Registry gives a repeatable, lineage-tracked, approval-gated workflow that reproduces any model version.

  • A. Automating a manual notebook does not add lineage, versioning or an approval gate; the same rollback problem persists.
  • B. Skipping retraining lets the model degrade as patterns change and does not solve the auditability requirement.
  • C. Versioned artifacts alone cannot reproduce a model because the data snapshot, code and hyperparameters are not captured.
  • D. Pipelines provides repeatable orchestration with lineage, and Model Registry adds versioning and approval before deployment.
Question 8Fundamentals of AI and ML

A social platform is adding image features and wants to know which of them a pre-trained computer vision service such as Amazon Rekognition can provide without the platform having to train any model of its own. Which TWO capabilities are available directly from the pre-trained service? (Select TWO.)

Choose 2.

  1. A.

    Recognising the platform's own internal product codes stamped on packaging, which appear in no public image dataset.

  2. B.

    Detecting unsafe or explicit content in an uploaded image so that it can be held for moderation.

  3. C.

    Detecting and comparing faces across images to check whether two photographs show the same person.

  4. D.

    Translating the captions that users write underneath their images into eleven other languages.

  5. E.

    Forecasting how many images the platform will receive next month from the upload history stored in Amazon S3.

Show answer

Answer: B, C

Pre-trained Amazon Rekognition covers content moderation and face detection and comparison out of the box.

  • A. Company-specific codes appear in no general dataset, so this needs training on the customer's own images.
  • B. Content moderation for unsafe or explicit imagery is a standard pre-trained Rekognition capability.
  • C. Face detection and face comparison are standard pre-trained capabilities of the service.
  • D. Translating text is a language task handled by Amazon Translate, not by a vision service.
  • E. Forecasting upload volume is a time-series problem on operational data, not a computer vision task.
Question 9Fundamentals of AI and ML

An energy utility has loaded five years of smart-meter readings for a consumption-forecasting project. Before cleaning the data or training anything, the team wants to carry out exploratory data analysis. Which TWO activities belong to the exploratory data analysis stage? (Select TWO.)

Choose 2.

  1. A.

    Run a hyperparameter search to select the number of training epochs that scores best.

  2. B.

    Register the approved model version so that it can be promoted to the production endpoint.

  3. C.

    Visualise consumption over time to identify seasonal patterns and obvious outliers.

  4. D.

    Summarise each column's distribution and count how many readings are missing or negative.

  5. E.

    Deploy the model to a real-time endpoint and enable request capture for later analysis.

Show answer

Answer: C, D

Summarising distributions and visualising patterns over time are exploratory data analysis activities.

  • A. A hyperparameter search is part of model training, which has not started yet.
  • B. Registering an approved version is a release-governance step near the end of the lifecycle.
  • C. Charting consumption over time to find seasonality and outliers is exploratory analysis.
  • D. Column summaries and counts of missing or impossible values are standard EDA outputs.
  • E. Deployment and request capture happen after a model exists and has passed evaluation.
Question 10Fundamentals of AI and ML

A logistics company trains a model that predicts how many minutes a delivery will be late, which is a continuous value. An analyst tries to report the model's accuracy and F1 score for it. Which type of metric should the analyst use instead?

  1. A.

    A classification metric such as precision, measuring how many flagged cases were genuinely positive

  2. B.

    An error metric such as mean absolute error, measuring how far predictions fall from actual values

  3. C.

    A classification metric such as recall, measuring how many actual positive cases the model found

  4. D.

    A business metric such as return on investment, measuring the financial value the model delivers

Show answer

Answer: B

A model predicting a continuous value is evaluated with error metrics, not classification metrics.

  • A. Precision is built from positive and negative counts and needs a discrete class outcome.
  • B. Continuous predictions are scored by error metrics such as mean absolute error.
  • C. Recall counts actual positives that were found, which presumes discrete classes a continuous prediction does not have.
  • D. ROI measures financial value delivered, not how close the predicted minutes were.

Keep going with 660 more AIF-C01 questions

Free papers every day, in the real exam formats, with progress by exam domain. Unlock every paper and timed mock exam when you are ready.