PMLE sample questions with answers

10 free practice questions for the Professional Machine Learning Engineer exam. Try each one, then open the answer to see why the right option wins and every other option loses.

Question 1Architecting low-code AI solutions

Calderon Legal Services stores 140,000 reviewed contract clauses in BigQuery, each with a house-style risk summary written by a paralegal. Zero-shot Gemini summaries are accurate but ignore the firm's fixed three-sentence format and its risk vocabulary, and prompt engineering has not fixed the formatting reliably. The data governance officer forbids copying clause text out of BigQuery, and the analytics team writes SQL rather than Python. The tuned behaviour must be callable from the nightly SQL job that already builds the review queue. What should you do?

  1. A.

    Generate embeddings for every clause with a BigQuery ML embedding model, then train a BigQuery ML logistic regression model on those embeddings to produce the house-style summaries.

  2. B.

    Create a BigQuery ML remote model over the base Gemini model and call AI.GENERATE_TEXT with a much longer system instruction that restates the three-sentence rule and lists the firm's risk vocabulary.

  3. C.

    Export the clause and summary pairs to Cloud Storage as JSONL files, and start a Gemini supervised tuning job from the Google Cloud console in the Agent Platform (formerly Vertex AI) project.

  4. D.

    Create a BigQuery ML remote model over a tunable Gemini model, with an AS SELECT clause that returns the clause as prompt and the paralegal summary as label, and call it with AI.GENERATE_TEXT.

Show answer

Answer: D

Supervised tuning of a Gemini remote model from BigQuery keeps the data in BigQuery and the workflow in SQL.

  • A. Logistic regression predicts a class label; it cannot generate a three-sentence summary.
  • B. More prompting is what has already failed to hold the format reliably; the scenario needs the behaviour learned, not restated.
  • C. Exporting clause text to Cloud Storage breaks the governance rule, and a console workflow moves the work away from the SQL team.
  • D. Supervised tuning from BigQuery ML runs from a CREATE MODEL statement over table data, and the tuned remote model is then used from SQL.
Question 2Architecting low-code AI solutions

An engineer at Tregothnan Shipping can call a BigQuery ML remote model over a base Gemini model on Gemini Enterprise Agent Platform (formerly Vertex AI) through a Cloud resource connection whose service account holds the Agent Platform User role. When she runs a CREATE MODEL statement with an AS SELECT clause to create a supervised-tuned version of the same Gemini model through that connection, the statement fails with a permission error on the Agent Platform side. The dataset, connection and training table are all in the same location. What should you do?

  1. A.

    Grant the Vertex AI Service Agent role to the connection's service account in the project where the tuned remote model is created, and then run the same CREATE MODEL statement again.

  2. B.

    Grant the engineer's own user account the Agent Platform User role in the project, because the tuning job runs with the credentials of the person who submits the CREATE MODEL statement.

  3. C.

    Grant the connection's service account the BigQuery Data Viewer role on the training table, because the tuning job reads the prompt and label columns through the connection's identity.

  4. D.

    Recreate the connection in the global location, because supervised tuning of a Gemini model is only supported through a connection that uses the global endpoint for its requests.

Show answer

Answer: A

Tuned remote models need the Vertex AI Service Agent role on the connection's service account; untuned ones need only Agent Platform User.

  • A. For a remote model that uses supervised tuning, the documentation requires the Vertex AI Service Agent role on the connection's service account; the Agent Platform User role suffices only for untuned models.
  • B. The remote model calls Agent Platform through the connection's service account, so the user's roles are not the missing piece.
  • C. The failure is on the Agent Platform side, and the connection already works for inference, so table read access is not the gap.
  • D. Location of the connection must match the dataset; nothing requires a global connection for tuning.
Question 3Architecting low-code AI solutions

An analyst at Draycott Water inherits a BigQuery ML boosted-tree model that predicts pipe bursts and needs to retrain it on a new quarter of data. Before changing anything, she wants a quick summary of the input features the model was trained on: minimum, maximum, mean and standard deviation for numeric columns, the number of categories for categorical columns, and how many NULL values each column had. She wants to get this from the model itself with one query. What should you do?

  1. A.

    Run a query that selects from ML.EVALUATE for the model without an input table, which returns summary statistics for every feature that the model saw during its original training run.

  2. B.

    Run a query that selects from ML.FEATURE_IMPORTANCE for the model, which returns the importance of each feature together with its minimum, maximum and count of missing values.

  3. C.

    Run a query that selects from ML.TRAINING_INFO for the model, which returns the loss and duration of each training iteration together with the distribution of every input feature column.

  4. D.

    Run a query on ML.FEATURE_INFO for the model, which returns each training input column with its minimum, maximum, mean, median, standard deviation, category count and null count.

Show answer

Answer: D

ML.FEATURE_INFO summarizes the input features a BigQuery ML model was trained on, including null and category counts.

  • A. ML.EVALUATE returns model quality metrics, not per-feature summary statistics.
  • B. ML.FEATURE_IMPORTANCE returns importance scores for tree models, not column statistics or null counts.
  • C. ML.TRAINING_INFO reports iteration-level training information such as loss and duration, not feature distributions.
  • D. ML.FEATURE_INFO returns statistics about each input feature used to train the model, including null_count and category_count.
Question 4Architecting low-code AI solutions

Pellucid Broadband wants to predict which one of five broadband plans each prospective customer will choose, so that the sales team can lead with it. An analyst trained five separate Agent Platform AutoML (formerly Vertex AI AutoML) binary classifiers, one per plan, each with the AUC PR objective, and then picks the plan with the highest score. Training five models every month is costly, and the five scores often conflict. What should you do?

  1. A.

    Keep the five binary classifiers but switch each of them to the maximize-au-roc objective, so that the scores produced by the five models become directly comparable with each other.

  2. B.

    Train one AutoML tabular regression model that predicts a plan number from 1 to 5, and round each predicted value to the nearest whole number to choose the plan for each prospect.

  3. C.

    Train one AutoML forecasting model that predicts the sequence of plan choices over time for each prospect, and take the first predicted value in the horizon as the recommended plan.

  4. D.

    Train one AutoML tabular classification model with the chosen plan as a five-class target, which frames the task as multi-class classification and uses the log loss objective for training.

Show answer

Answer: D

Predicting one of several discrete options is multi-class classification: one AutoML model, log loss objective.

  • A. Separate binary models still disagree and still cost five trainings; changing the objective does not calibrate them against each other.
  • B. Plan numbers are categories, not quantities, so regression imposes a false order between plans.
  • C. Forecasting predicts a sequence of values over time for a series, which is not this problem.
  • D. Choosing one of three or more discrete classes is multi-class classification; one model replaces five, and log loss is the only supported objective for it.
Question 5Architecting low-code AI solutions

Kestrel Logistics extracts fields from 9,000 scanned supplier invoices a month with a Document AI Invoice Parser processor. Two years ago a developer hard-coded a specific processor version ID in the process request URL. Last week Google deprecated that stable version and every request began to fail. The ML lead wants version retirements never to cause an outage again, but also wants to test any new stable version on the team's own labelled invoices before production traffic moves to it. What should you do?

  1. A.

    Call the processor endpoint without a version ID so the default version is used, and make a newer stable version the default only after evaluating it on labelled invoices.

  2. B.

    Hard-code the newest stable version ID in the request URL instead, and create a new processor for each future version so that old and new versions never share an endpoint.

  3. C.

    Point the request URL at the stable channel, so that each request is served by the newest stable version as soon as Google releases it, without any change by the team.

  4. D.

    Point the request URL at the rc channel, so that the processor always uses the most recent release candidate and never depends on a stable version that Google can later deprecate.

Show answer

Answer: A

Call the processor without a version ID and manage the default version; a deprecated default moves to the latest stable version, while pinned IDs fail once deprecated.

  • A. The default version survives deprecations by moving to the latest stable version, and the team can evaluate a new version before setting it as the default.
  • B. A pinned version ID fails again when that version is deprecated, and a processor per version multiplies endpoints without solving the lifecycle.
  • C. The stable channel uses each new stable version as soon as it is released, so production traffic moves before the team has tested it.
  • D. Release candidates are experimental versions that are not production-quality, and they change regularly without the team's testing.
Question 6Architecting low-code AI solutions

Oldcastle Energy is building a retrieval-augmented assistant over 15,000 engineering reports in PDF, DOCX and PPTX formats. A proof of concept that ran plain OCR and split the text every 500 characters returns poor answers: tables are flattened into unreadable text, and chunks cut across headings so retrieved passages lose their context. The team wants documents parsed with their structure intact and chunked for retrieval. What should you do?

  1. A.

    Convert every report to images and send them to the Cloud Vision API TEXT_DETECTION feature, so that the text is recognized page by page before the pages are chunked for retrieval.

  2. B.

    Process the reports with the Document AI Layout Parser, which recognizes headings, tables and lists and produces context-aware chunks, and index those chunks for the retrieval step.

  3. C.

    Keep the plain OCR step but reduce the chunk size to 200 characters, so that each chunk is less likely to cross a heading and more of the retrieved passages stay on one topic.

  4. D.

    Process the reports with the Document AI Form Parser, and index each key-value pair that it returns from the engineering reports as a separate retrieval chunk for the assistant to search.

Show answer

Answer: B

For RAG over structured documents, parse and chunk with the Document AI Layout Parser rather than plain OCR and fixed-size splits.

  • A. Sparse-text OCR on page images discards structure and native text, the very problem being solved.
  • B. The Layout Parser extracts text, tables and lists from PDF, DOCX, PPTX and other formats and creates context-aware chunks for search and RAG.
  • C. Smaller fixed chunks still ignore structure and fragment tables further.
  • D. Form Parser targets key-value pairs in forms; engineering reports are long narrative documents, and key-value chunks lose context.
Question 7Architecting low-code AI solutions

Arkwright Tools translates product manuals from English into eight languages with the Cloud Translation API. Reviewers keep finding that brand names such as the TorqueMax product line are translated literally, and that the word 'chuck' is rendered as a verb instead of the drill part. The company has a terminology list of about 600 source and target term pairs and wants every translation to respect it consistently. What should you do?

  1. A.

    Create a glossary from the terminology list with Cloud Translation - Advanced, store the glossary file in Cloud Storage, and reference the glossary in every translation request for the manuals.

  2. B.

    Post-process every translated manual with a script that searches for incorrect renderings of each of the 600 terms and replaces them with the approved translation for the target language.

  3. C.

    Keep using Cloud Translation - Basic but wrap the brand names in quotation marks before each request, because quoted text is never translated and the drill part will then be inferred correctly.

  4. D.

    Train a custom NMT model on the 600 term pairs alone, because a custom model trained on the terminology list learns to apply those terms consistently throughout full manual sentences.

Show answer

Answer: A

Consistent handling of product names and ambiguous domain terms is a Cloud Translation - Advanced glossary.

  • A. A glossary is the documented way to make Cloud Translation consistently translate domain terms, product names and ambiguous words.
  • B. Search-and-replace cannot handle inflection and context in eight languages and is brittle to maintain.
  • C. Quotation marks are not a documented way to protect terms, and glossaries are a feature of the Advanced edition.
  • D. Custom NMT training is a heavier project that adapts the whole model from segment pairs; it does not guarantee each listed term is always applied.
Question 8Architecting low-code AI solutions

Kelsall Software wants its help centre to return the most relevant of 30,000 support articles when a customer types a question, and a retrieval step will later feed a generative answer. A prototype sends every article and the question to a large Gemini model and asks it to rate relevance, which is slow and expensive. The team is choosing a different model from Model Garden on Gemini Enterprise Agent Platform (formerly Vertex AI) for this retrieval step. What should you do?

  1. A.

    Use a text embedding model, embedding articles with the RETRIEVALDOCUMENT task type and customer questions with RETRIEVALQUERY, and rank articles by vector similarity to the question.

  2. B.

    Use a smaller Gemini Flash-Lite model to rate the relevance of every article for each question, because a cheaper generative model makes the comparison of all 30,000 articles fast enough.

  3. C.

    Use a Gemini image generation model to render every article as a thumbnail, and compare the thumbnails with a rendering of the question to find the visually closest support article.

  4. D.

    Use the Cloud Translation API to translate every question and article into English first, and then rank the articles by the number of words that they share with the translated question.

Show answer

Answer: A

For semantic retrieval, select an embedding model and use query and document task types, not a generative model.

  • A. Embedding models turn text into vectors for semantic search, and retrieval task types optimize document and query embeddings for question-to-answer matching.
  • B. Any generative model scoring every article per question scales with the corpus size and stays slow and costly.
  • C. Rendering text as images has nothing to do with semantic relevance.
  • D. Word overlap after translation is lexical matching, which misses semantically related articles.
Question 9Architecting low-code AI solutions

Clearwell Mortgages receives each application as a single 60-page PDF that combines an application form, pay slips, bank statements and a photo ID in varying order. Each document type must go to a different Document AI extraction processor, and today a clerk splits the packets by hand before upload. The team wants the packets separated and each part labelled automatically before extraction. What should you do?

  1. A.

    Send each packet directly to the Document AI Bank Statement Parser processor, because it extracts the entities of the first supported document it finds and ignores all of the other pages in the file.

  2. B.

    Send each packet to the Document AI Enterprise Document OCR processor, and split the file wherever the page text starts with a new heading so that each part can be routed afterwards.

  3. C.

    Send each packet to a Document AI Custom Splitter processor trained on the document classes, and route every identified logical document to the matching extraction processor by its class.

  4. D.

    Send each packet to a Document AI Custom Classifier processor, and forward the entire 60-page file to the extraction processor that matches the single class returned for the packet.

Show answer

Answer: C

Composite PDFs with several document types need a Document AI Custom Splitter before type-specific extraction.

  • A. A single parser handles only one document type and ignores the rest of the packet.
  • B. Heading-based rules are fragile and do not classify the resulting parts.
  • C. A Custom Splitter identifies the pages that make up each logical document in a composite file and classifies them, so each part can go to the right extractor.
  • D. A classifier assigns one class to the whole file; it does not find the page boundaries of the documents inside a composite packet.
Question 10Architecting low-code AI solutions

Northbay Telecom trains an Agent Platform AutoML (formerly Vertex AI AutoML) tabular classifier that predicts whether a subscriber will churn in the next 30 days. Offline the model reports an AUC PR of 0.93, but when the campaign team used it on last month's subscribers the lift almost vanished. The dataset covers 24 months, each row is one subscriber-month, and the default random data split was used. Churn behaviour shifted noticeably after a tariff change nine months ago. You must make the offline metric predictive of live performance. What should you do?

  1. A.

    Increase the AutoML training budget in node hours, so that the architecture search explores more candidate models and has a better chance of finding one that generalizes to recent months.

  2. B.

    Configure a manual data split that assigns 80 percent of subscriber IDs to training and 20 percent to test, so that no individual subscriber appears in both the training and test sets.

  3. C.

    Switch the optimization objective from AUC PR to log loss, so that the predicted probabilities are better calibrated and the campaign team can rely on the scores they receive.

  4. D.

    Configure a chronological data split on the subscriber-month timestamp column, so that the training, validation and test sets cover successive time ranges and the test set holds the most recent months.

Show answer

Answer: D

When behaviour drifts over time, evaluate the AutoML model on the most recent period with a chronological split.

  • A. A bigger search optimizes the same leaky evaluation; it cannot fix a test set that does not resemble production.
  • B. Splitting by subscriber removes one leak but still mixes months, so the test set does not represent the post-change period the model will face.
  • C. Calibration changes how probabilities are scored, not which rows are in the test set, so the metric stays optimistic.
  • D. A chronological split evaluates on the newest data, which mirrors how the model is used and exposes the post-tariff shift the random split hid.

Keep going with 493 more PMLE questions

Free papers every day, in the real exam formats, with progress by exam domain. Unlock every paper and timed mock exam when you are ready.

PMLE sample questions with answers (10 free) · CertifyCloudx