Calderon Legal Services stores 140,000 reviewed contract clauses in BigQuery, each with a house-style risk summary written by a paralegal. Zero-shot Gemini summaries are accurate but ignore the firm's fixed three-sentence format and its risk vocabulary, and prompt engineering has not fixed the formatting reliably. The data governance officer forbids copying clause text out of BigQuery, and the analytics team writes SQL rather than Python. The tuned behaviour must be callable from the nightly SQL job that already builds the review queue. What should you do?
- A.
Generate embeddings for every clause with a BigQuery ML embedding model, then train a BigQuery ML logistic regression model on those embeddings to produce the house-style summaries.
- B.
Create a BigQuery ML remote model over the base Gemini model and call AI.GENERATE_TEXT with a much longer system instruction that restates the three-sentence rule and lists the firm's risk vocabulary.
- C.
Export the clause and summary pairs to Cloud Storage as JSONL files, and start a Gemini supervised tuning job from the Google Cloud console in the Agent Platform (formerly Vertex AI) project.
- D.
Create a BigQuery ML remote model over a tunable Gemini model, with an AS SELECT clause that returns the clause as prompt and the paralegal summary as label, and call it with AI.GENERATE_TEXT.
Show answer
Answer: D
Supervised tuning of a Gemini remote model from BigQuery keeps the data in BigQuery and the workflow in SQL.
- A. Logistic regression predicts a class label; it cannot generate a three-sentence summary.
- B. More prompting is what has already failed to hold the format reliably; the scenario needs the behaviour learned, not restated.
- C. Exporting clause text to Cloud Storage breaks the governance rule, and a console workflow moves the work away from the SQL team.
- D. Supervised tuning from BigQuery ML runs from a CREATE MODEL statement over table data, and the tuned remote model is then used from SQL.
