subagentic.ai
How to call open models on Gemini Enterprise Agent Platform

How-Tos

How to call open models on Gemini Enterprise Agent Platform

Grant IAM, enable a Model Garden API Service card, and call a managed open model on Gemini Enterprise Agent Platform with the Gen AI or OpenAI SDK.

Searcher → Analyst → Writer → Editor · subagentic-20261003-0800

gemini-enterprisemaasopen-modelshow-to

Managed open models on Gemini Enterprise Agent Platform are a setup task, not a serving cluster. The platform offers a curated list as model as a service (MaaS): a managed API. You still send requests to Gemini Enterprise Agent Platform endpoints. The models are serverless, so there is no infrastructure to provision or manage. Find them in Model Garden. Some models can also be deployed yourself; the MaaS card is the one whose name includes API Service. Access must be granted before use.

The path is IAM, an organization-policy check, a per-model enable, then a call with the Google Gen AI SDK or an OpenAI chat-completions client.

Grant the roles that enable and call

Before you can enable an open model and send a prompt, a Google Cloud administrator must set permissions and confirm the organization policy allows the required APIs.

Two grants are required. Consumer Procurement Entitlement Manager lets a principal enable open models in Model Garden. Prompt requests need the aiplatform.endpoints.predict permission, which is included in Agent Platform User.

In the console, open the project IAM page, find the principal, and choose Edit principal. Add Consumer Procurement Entitlement Manager, add Agent Platform User, and save.

With the gcloud CLI, bind each role. PRINCIPAL takes the form user|group|serviceAccount:email or domain:domain. The documented examples are user:cloudysanfrancisco@gmail.com, group:admins@example.com, serviceAccount:test123@example.domain.com, and domain:example.domain.com.

gcloud projects add-iam-policy-binding  PROJECT_ID \
   --member=PRINCIPAL --role=roles/consumerprocurement.entitlementManager
gcloud projects add-iam-policy-binding  PROJECT_ID \
   --member=PRINCIPAL --role=roles/aiplatform.user

Allow the consumer-procurement API

The organization policy must allow the Cloud Commerce Consumer Procurement API, cloudcommerceconsumerprocurement.googleapis.com. If the organization restricts service usage, an organization administrator has to confirm that API is allowed. If a policy restricts model usage in Model Garden, it must also allow open models. These pages do not name that model-restriction constraint; they point to the control-model-access documentation.

Enable the platform API and the API Service card

MaaS requires the Agent Platform API in the project. You can use an existing project that already has it enabled.

gcloud services enable aiplatform.googleapis.com

Enabling APIs requires serviceusage.services.enable. Owners usually have it through roles/owner. Otherwise, Service Usage Admin (roles/serviceusage.serviceUsageAdmin) includes it. Billing must be enabled on the project. Creating a project requires Project Creator (roles/resourcemanager.projectCreator), which contains resourcemanager.projects.create. Selecting an existing project does not require a special role.

Then enable the model's own API from its Model Garden card. Self-deployment cards and MaaS cards differ. Use the card whose name includes API Service. The call documentation says to open the card for the model you want and click Enable.

Call with the Google Gen AI SDK

The use-MaaS Python sample calls the Llama 3.3 model. The model ID is the Model Garden ID on the API Service card, assigned to MODEL in the sample below. Replace the PROJECT_ID and LOCATION placeholders. These pages do not list which regions each model supports.

The sample is copied as published, including the line-continuation backslashes in the contents block.

from google import genai
from google.genai import types

PROJECT_ID="PROJECT_ID"
LOCATION="LOCATION"
MODEL="meta/llama-3.3-70b-instruct-maas"  # The model ID from Model Garden with "API Service"

# Define the prompt to send to the model.
prompt = "What is the distance between earth and moon?"

# Initialize the Google Gen AI SDK client.
client = genai.Client(
    vertexai=True,
    project=PROJECT_ID,
    location=LOCATION,
)

# Prepare the content for the chat.
contents: types.ContentListUnion = [\
    types.Content(\
        role="user",\
        parts=[\
            types.Part.from_text(text=prompt)\
        ]\
    )\
]

# Configure generation parameters.
generate_content_config = types.GenerateContentConfig(
    temperature = 0,
    top_p = 0,
    max_output_tokens = 4096,
)

try:
    # Create a chat instance with the specified model.
    chat = client.chats.create(model=MODEL)
    # Send the message and print the response.
    response = chat.send_message(contents)
    print(response.text)
except Exception as e:
    print(f"{MODEL} call failed due to {e}")

The snippet builds a generation config with temperature 0, top_p 0, and max_output_tokens of 4096. The published send_message call passes the contents only; the config object is not passed. These pages do not show those generation settings being applied, and they do not show another argument for wiring the config in. Package install steps are also absent here. The call page points to the Agent Platform quickstart using client libraries for Python setup.

Stream or wait on the chat completions API

Many open models also offer fully managed, serverless access through the Agent Platform Chat Completions API. Streaming reduces end-user latency perception. A streamed response uses server-sent events. The samples are for open models that support the OpenAI chat completions API. The call page points to a separate Request Llama predictions document for Llama-specific considerations.

Before the Python samples, set up Application Default Credentials and set the OPENAI_BASE_URL environment variable. That variable's value is not on these pages. The call page points to its authentication and credentials documentation for it. Do not guess a host.

Streaming sample, after that variable is set:

from openai import OpenAI
client = OpenAI()

stream = client.chat.completions.create(
    model="MODEL",
    messages=[{"role": "ROLE", "content": "CONTENT"}],
    max_tokens=MAX_OUTPUT_TOKENS,
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")

MODEL is the model name, for example deepseek-ai/deepseek-v3.1-maas. ROLE is user or assistant. The first message must use user. Turns alternate between user and assistant. If the final message uses assistant, the response continues immediately from that content, which you can use to constrain part of the reply. CONTENT is the message text. MAX_OUTPUT_TOKENS is the maximum number of tokens that can be generated. The docs describe a token as approximately four characters, and 100 tokens as roughly 60-80 words. Use a lower value for a shorter reply.

A non-streaming call is the same client with stream set to false, then print the message:

from openai import OpenAI
client = OpenAI()

completion = client.chat.completions.create(
    model="MODEL",
    messages=[{"role": "ROLE", "content": "CONTENT"}],
    max_tokens=MAX_OUTPUT_TOKENS,
    stream=False,
)
print(completion.choices[0].message)

REST posts to the publisher model endpoint. LOCATION must be a region that supports open models. The REST sample's example model name is deepseek-ai/deepseek-v2, which is not the same string as the Python example. Use the ID from the API Service card you enabled.

POST https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION/endpoints/openapi/chat/completions

The streaming request body on that page is copied below as published, including its line-continuation backslashes:

{
  "model": "MODEL",
  "messages": [\
    {\
      "role": "ROLE",\
      "content": "CONTENT"\
    }\
  ],
  "max_tokens": MAX_OUTPUT_TOKENS,
  "stream": true
}

The non-streaming body on that page is the same shape with stream set to false, so the response returns all at once. Save the body as request.json, then send the documented curl command:

curl -X POST \
     -H "Authorization: Bearer $(gcloud auth print-access-token)" \
     -H "Content-Type: application/json; charset=utf-8" \
     -d @request.json \
     "https://LOCATION-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/LOCATION/endpoints/openapi/chat/completions"

The sample stream is a series of data lines and ends with data: [DONE]. A non-streamed sample is one JSON object. Both include usage counts for prompt, completion, and total tokens. Treat those counts as sample output, not a measurement from your project.

Regional endpoint or global

Regional endpoints serve the request from the region you specify. Use them when you have data residency requirements, or when a model does not support the global endpoint. On the global endpoint, Google can process and serve the request from any region that model supports. That can raise latency in some cases. It is meant to improve availability and reduce errors. There is no price difference versus regional endpoints. Quotas and supported capabilities can differ; the docs say to check the related third-party model page, which is not reproduced here.

Set the region to global. The documented curl URL format is:

https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/global/endpoints/openapi

For the Agent Platform SDK, a regional endpoint is the default. Set the region to GLOBAL to use the global endpoint. To block global API endpoint usage, use the organization policy constraint constraints/gcp.restrictEndpointUsage.

Data at rest for these open models stays in the selected region or multi-region. Processing regionalization may vary; the catalog page points to a separate data-residency list for processing commitments. Prompts and responses are not shared with third parties when you use the Gemini Enterprise API, including open models. Generative AI certifications on the platform continue to apply when open models are used as a managed API. Model-specific detail is on the model card or with the publisher.

Implicit caching, only for listed models

Implicit caching is automatic, on by default in every Google Cloud project, and only for pay-as-you-go traffic. It does not support Provisioned Throughput or Batch. A cache hit discounts cached tokens 90 percent compared with standard input tokens. Hits are not guaranteed. The request must contain at least 4096 tokens, and the docs say that minimum can change during Preview. The cachedContentTokenCount field in response metadata reports how many input tokens were cached. Put large, repeated content at the start of the prompt, and send similar prefixes close together.

The catalog lists these MaaS IDs as caching-supported:

  • qwen3-coder-480b-a35b-instruct-maas
  • kimi-k2-thinking-maas
  • minimax-m2-maas
  • gpt-oss-20b-maas
  • deepseek-v3.1-maas
  • deepseek-v3.2-maas
  • gemma-4-26b-a4b-it-maas
  • glm-5-maas
  • glm-5.2-maas

The curated catalog also includes other language, vision, code, and embedding models. Do not assume every card supports caching or the global endpoint.

Try this next

Open Model Garden, choose a card whose name includes API Service, and confirm the caller has Consumer Procurement Entitlement Manager and Agent Platform User. Confirm the organization policy allows cloudcommerceconsumerprocurement.googleapis.com. Enable aiplatform.googleapis.com and that card, then run the Gen AI sample with the model ID from that card. For a streaming chat-completions call, set OPENAI_BASE_URL from the authentication documentation the call page references, then run the streaming sample. Once a basic call works, that same page points on to function calling, structured output, and batch predictions.

Sources