Skip to Content
接口调用Google Vertex AI API

Google Vertex AI API

本文按 Vertex AI 官方文档整理,完整展示 Vertex AI 上 Gemini 原生接口、Gemini OpenAI-compatible 入口、Claude partner model 入口、以及 OpenAI 风格 managed API 入口的请求地址、请求参数、请求示例、请求返回结果、请求返回示例。

鉴权方式

  • 使用 Google Cloud Access Token。
  • 常见 Header:Authorization: Bearer $(gcloud auth print-access-token)

请求地址

1. Vertex AI 上的 Gemini 原生接口

  • 方法:POST
  • 地址:https://aiplatform.googleapis.com/v1/projects/{project}/locations/{location}/publishers/google/models/{model}:generateContent

2. Vertex AI 上的 Gemini OpenAI-compatible 入口

  • 方法:POST
  • 地址:https://aiplatform.googleapis.com/v1/projects/{project}/locations/{location}/endpoints/openapi/chat/completions

3. Vertex AI 上的 Claude partner model 原生入口

  • 方法:POST
  • 地址:https://{location}-aiplatform.googleapis.com/v1/projects/{project}/locations/{location}/publishers/anthropic/models/{model}:rawPredict

4. Vertex AI 上的 OpenAI 风格 managed API 入口

  • 方法:POST
  • 地址:https://{location}-aiplatform.googleapis.com/v1/projects/{project}/locations/{location}/endpoints/openapi/chat/completions

请求参数

Gemini 原生接口请求参数

参数类型必填说明
contentsarray是对话内容。
contents[].rolestring否user / model。
contents[].partsarray是文本、图片、文件等内容块。
systemInstructionobject否系统提示。
toolsarray否工具定义。
toolConfigobject否工具配置。
safetySettingsarray否安全设置。
generationConfigobject否生成配置,如 temperature、topP、topK、maxOutputTokens。
cachedContentstring否缓存上下文引用。

Gemini OpenAI-compatible 入口请求参数

参数类型必填说明
modelstring是例如 google/gemini-2.0-flash-001。
messagesarray是OpenAI 风格消息数组。
messages[].rolestring是system、user、assistant 等。
messages[].contentstring | array是文本或多模态内容。
temperaturenumber否采样温度。
top_pnumber否nucleus sampling。
max_tokensinteger否最大输出 token 数。
streamboolean否是否流式输出。
toolsarray否工具定义。
tool_choicestring | object否工具选择策略。
response_formatobject否输出格式。
reasoning_effortstring否推理强度。
extra_body.google.thinking_configobject否Gemini 特定 thinking 配置。

Claude partner model 请求参数

参数类型必填说明
anthropic_versionstring是Anthropic 协议版本。
messagesarray是Claude 消息数组。
max_tokensinteger是最大输出 token 数。
systemstring | array否系统提示。
temperaturenumber否采样温度。
top_pnumber否nucleus sampling。
top_kinteger否候选 token 上限。
toolsarray否工具定义。
tool_choicestring | object否工具选择策略。

OpenAI 风格 managed API 请求参数

参数类型必填说明
modelstring是例如 gpt-oss-120b-maas 等 Vertex 托管模型 ID。
messagesarray是OpenAI 风格消息数组。
temperaturenumber否采样温度。
top_pnumber否nucleus sampling。
max_tokensinteger否最大输出 token 数。
streamboolean否是否流式输出。
toolsarray否工具定义。

请求示例

Gemini 原生接口

curl -X POST \ -H "Authorization: Bearer $(gcloud auth print-access-token)" \ -H "Content-Type: application/json" \ "https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/us-central1/publishers/google/models/gemini-2.5-flash:generateContent" \ -d '{ "contents": [ { "role": "user", "parts": [ { "text": "Explain GGPU Platform in one paragraph." } ] } ], "generationConfig": { "temperature": 0.7, "maxOutputTokens": 256 } }'

Gemini OpenAI-compatible 入口

curl -X POST \ -H "Authorization: Bearer $(gcloud auth print-access-token)" \ -H "Content-Type: application/json" \ "https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/us-central1/endpoints/openapi/chat/completions" \ -d '{ "model": "google/gemini-2.0-flash-001", "messages": [ { "role": "user", "content": "Explain GGPU Platform in one paragraph." } ], "temperature": 0.7, "max_tokens": 256 }'

Claude partner model 入口

curl -X POST \ -H "Authorization: Bearer $(gcloud auth print-access-token)" \ -H "Content-Type: application/json" \ "https://us-east5-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/us-east5/publishers/anthropic/models/claude-sonnet-4-5@20250929:rawPredict" \ -d '{ "anthropic_version": "vertex-2023-10-16", "max_tokens": 256, "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Explain GGPU Platform in one paragraph." } ] } ], "temperature": 0.7 }'

OpenAI 风格 managed API 入口

curl -X POST \ -H "Authorization: Bearer $(gcloud auth print-access-token)" \ -H "Content-Type: application/json" \ "https://us-central1-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/us-central1/endpoints/openapi/chat/completions" \ -d '{ "model": "gpt-oss-120b-maas", "messages": [ { "role": "user", "content": "Explain GGPU Platform in one paragraph." } ], "temperature": 0.7, "max_tokens": 256 }'

Omni HOT

Vertex AI 的 Interactions 视频生成目前暂不支持 gemini-omni-1.1-flash,请使用 vertex-ai/gemini-omni-flash-preview。以下示例使用平台创建的访问令牌,并将令牌放在 GGPU_API_KEY 环境变量中。

异步生成

设置 "background": true 后,接口会立即创建后台任务并返回平台任务信息。请保存返回的 id,再使用统一任务接口获取状态和结果。

curl --location --request POST "https://api.ggpu.ai/v1beta/interactions" \ --header "x-goog-api-key: $GGPU_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "vertex-ai/gemini-omni-flash-preview", "input": "小狗跳舞", "response_format": { "type": "video", "aspect_ratio": "16:9" }, "generation_config": { "video_config": { "task": "text_to_video" } }, "background": true, "store": true }'

使用返回的任务 ID 查询异步任务状态:

curl --location --request GET "https://api.ggpu.ai/v1/tasks/<task_id>" \ --header "Authorization: Bearer $GGPU_API_KEY"

同步生成

省略 background(或显式设置为 false)时,接口按同步 interaction 处理,请求会等待生成结果返回。

curl --location --request POST "https://api.ggpu.ai/v1beta/interactions" \ --header "x-goog-api-key: $GGPU_API_KEY" \ --header "Content-Type: application/json" \ --data-raw '{ "model": "vertex-ai/gemini-omni-flash-preview", "input": "小狗跳舞", "response_format": { "type": "video", "aspect_ratio": "16:9" }, "generation_config": { "video_config": { "task": "text_to_video" } }, "store": true }'

请求返回结果

Gemini 原生接口返回结果

参数类型说明
candidatesarray候选输出。
candidates[].contentobject生成内容。
candidates[].finishReasonstring停止原因。
promptFeedbackobjectPrompt 反馈。
usageMetadataobjecttoken 用量。
modelVersionstring实际模型版本。

Gemini / OpenAI 风格 chat completions 返回结果

参数类型说明
idstringcompletion ID。
objectstring对象类型。
createdinteger创建时间。
modelstring实际模型。
choicesarraycompletion 候选。
choices[].messageobjectassistant 消息。
choices[].finish_reasonstring停止原因。
usageobjecttoken 用量。

Claude partner model 返回结果

参数类型说明
idstring消息 ID。
typestring通常为 message。
rolestring通常为 assistant。
contentarray输出内容块。
modelstring实际模型。
stop_reasonstring停止原因。
usageobjecttoken 用量。

请求返回示例

Gemini 原生接口返回示例

{ "candidates": [ { "content": { "role": "model", "parts": [ { "text": "GGPU Platform helps teams design, execute, and scale AI workflows on top of Vertex AI and external model providers." } ] }, "finishReason": "STOP" } ], "usageMetadata": { "promptTokenCount": 13, "candidatesTokenCount": 21, "totalTokenCount": 34 }, "modelVersion": "gemini-2.5-flash" }

Gemini / OpenAI 风格 chat completions 返回示例

{ "id": "chatcmpl-vertex-01", "object": "chat.completion", "created": 1741572201, "model": "google/gemini-2.0-flash-001", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "GGPU Platform helps teams design, execute, and scale AI workflows on top of Vertex AI and external model providers." }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 13, "completion_tokens": 20, "total_tokens": 33 } }

Claude partner model 返回示例

{ "id": "msg_vertex_01", "type": "message", "role": "assistant", "model": "claude-sonnet-4-5@20250929", "content": [ { "type": "text", "text": "GGPU Platform helps teams design, execute, and scale AI workflows on top of Vertex AI and external model providers." } ], "stop_reason": "end_turn", "stop_sequence": null, "usage": { "input_tokens": 15, "output_tokens": 20 } }

说明

  • Vertex AI 的 Omni Interactions 视频生成暂不支持 gemini-omni-1.1-flash,请使用 vertex-ai/gemini-omni-flash-preview。
  • background: true 表示异步任务,创建后使用 /v1/tasks/<task_id> 轮询;省略或设为 false 时按同步 interaction 处理。
  • Vertex AI 不是单一的一套接口,Gemini、Claude partner models、OpenAI-compatible 入口各自有不同 schema。
  • 如果文档目标是“官方原生展示”,应按入口分别呈现,而不是合并成单一最小示例。
  • 更细的 streaming、function/tool calling、多模态 part 类型,请以 Vertex AI 官方文档为准。