Google Vertex AI API
本文按 Vertex AI 官方文档整理,完整展示 Vertex AI 上 Gemini 原生接口、Gemini OpenAI-compatible 入口、Claude partner model 入口、以及 OpenAI 风格 managed API 入口的请求地址、请求参数、请求示例、请求返回结果、请求返回示例。
鉴权方式
- 使用 Google Cloud Access Token。
- 常见 Header:
Authorization: Bearer $(gcloud auth print-access-token)
请求地址
1. Vertex AI 上的 Gemini 原生接口
- 方法:
POST - 地址:
https://aiplatform.googleapis.com/v1/projects/{project}/locations/{location}/publishers/google/models/{model}:generateContent
2. Vertex AI 上的 Gemini OpenAI-compatible 入口
- 方法:
POST - 地址:
https://aiplatform.googleapis.com/v1/projects/{project}/locations/{location}/endpoints/openapi/chat/completions
3. Vertex AI 上的 Claude partner model 原生入口
- 方法:
POST - 地址:
https://{location}-aiplatform.googleapis.com/v1/projects/{project}/locations/{location}/publishers/anthropic/models/{model}:rawPredict
4. Vertex AI 上的 OpenAI 风格 managed API 入口
- 方法:
POST - 地址:
https://{location}-aiplatform.googleapis.com/v1/projects/{project}/locations/{location}/endpoints/openapi/chat/completions
请求参数
Gemini 原生接口请求参数
Gemini OpenAI-compatible 入口请求参数
Claude partner model 请求参数
OpenAI 风格 managed API 请求参数
请求示例
Gemini 原生接口
curl
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
"https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/us-central1/publishers/google/models/gemini-2.5-flash:generateContent" \
-d '{
"contents": [
{
"role": "user",
"parts": [
{
"text": "Explain GGPU Platform in one paragraph."
}
]
}
],
"generationConfig": {
"temperature": 0.7,
"maxOutputTokens": 256
}
}'Gemini OpenAI-compatible 入口
curl
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
"https://aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/us-central1/endpoints/openapi/chat/completions" \
-d '{
"model": "google/gemini-2.0-flash-001",
"messages": [
{
"role": "user",
"content": "Explain GGPU Platform in one paragraph."
}
],
"temperature": 0.7,
"max_tokens": 256
}'Claude partner model 入口
curl
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
"https://us-east5-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/us-east5/publishers/anthropic/models/claude-sonnet-4-5@20250929:rawPredict" \
-d '{
"anthropic_version": "vertex-2023-10-16",
"max_tokens": 256,
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "Explain GGPU Platform in one paragraph."
}
]
}
],
"temperature": 0.7
}'OpenAI 风格 managed API 入口
curl
curl -X POST \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "Content-Type: application/json" \
"https://us-central1-aiplatform.googleapis.com/v1/projects/PROJECT_ID/locations/us-central1/endpoints/openapi/chat/completions" \
-d '{
"model": "gpt-oss-120b-maas",
"messages": [
{
"role": "user",
"content": "Explain GGPU Platform in one paragraph."
}
],
"temperature": 0.7,
"max_tokens": 256
}'Omni HOT
Vertex AI 的 Interactions 视频生成目前暂不支持 gemini-omni-1.1-flash,请使用 vertex-ai/gemini-omni-flash-preview。以下示例使用平台创建的访问令牌,并将令牌放在 GGPU_API_KEY 环境变量中。
异步生成
设置 "background": true 后,接口会立即创建后台任务并返回平台任务信息。请保存返回的 id,再使用统一任务接口获取状态和结果。
curl --location --request POST "https://api.ggpu.ai/v1beta/interactions" \
--header "x-goog-api-key: $GGPU_API_KEY" \
--header "Content-Type: application/json" \
--data-raw '{
"model": "vertex-ai/gemini-omni-flash-preview",
"input": "小狗跳舞",
"response_format": {
"type": "video",
"aspect_ratio": "16:9"
},
"generation_config": {
"video_config": {
"task": "text_to_video"
}
},
"background": true,
"store": true
}'使用返回的任务 ID 查询异步任务状态:
curl --location --request GET "https://api.ggpu.ai/v1/tasks/<task_id>" \
--header "Authorization: Bearer $GGPU_API_KEY"同步生成
省略 background(或显式设置为 false)时,接口按同步 interaction 处理,请求会等待生成结果返回。
curl --location --request POST "https://api.ggpu.ai/v1beta/interactions" \
--header "x-goog-api-key: $GGPU_API_KEY" \
--header "Content-Type: application/json" \
--data-raw '{
"model": "vertex-ai/gemini-omni-flash-preview",
"input": "小狗跳舞",
"response_format": {
"type": "video",
"aspect_ratio": "16:9"
},
"generation_config": {
"video_config": {
"task": "text_to_video"
}
},
"store": true
}'请求返回结果
Gemini 原生接口返回结果
Gemini / OpenAI 风格 chat completions 返回结果
Claude partner model 返回结果
请求返回示例
Gemini 原生接口返回示例
{
"candidates": [
{
"content": {
"role": "model",
"parts": [
{
"text": "GGPU Platform helps teams design, execute, and scale AI workflows on top of Vertex AI and external model providers."
}
]
},
"finishReason": "STOP"
}
],
"usageMetadata": {
"promptTokenCount": 13,
"candidatesTokenCount": 21,
"totalTokenCount": 34
},
"modelVersion": "gemini-2.5-flash"
}Gemini / OpenAI 风格 chat completions 返回示例
{
"id": "chatcmpl-vertex-01",
"object": "chat.completion",
"created": 1741572201,
"model": "google/gemini-2.0-flash-001",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"content": "GGPU Platform helps teams design, execute, and scale AI workflows on top of Vertex AI and external model providers."
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 13,
"completion_tokens": 20,
"total_tokens": 33
}
}Claude partner model 返回示例
{
"id": "msg_vertex_01",
"type": "message",
"role": "assistant",
"model": "claude-sonnet-4-5@20250929",
"content": [
{
"type": "text",
"text": "GGPU Platform helps teams design, execute, and scale AI workflows on top of Vertex AI and external model providers."
}
],
"stop_reason": "end_turn",
"stop_sequence": null,
"usage": {
"input_tokens": 15,
"output_tokens": 20
}
}说明
- Vertex AI 的 Omni Interactions 视频生成暂不支持
gemini-omni-1.1-flash,请使用vertex-ai/gemini-omni-flash-preview。 background: true表示异步任务,创建后使用/v1/tasks/<task_id>轮询;省略或设为false时按同步 interaction 处理。- Vertex AI 不是单一的一套接口,Gemini、Claude partner models、OpenAI-compatible 入口各自有不同 schema。
- 如果文档目标是“官方原生展示”,应按入口分别呈现,而不是合并成单一最小示例。
- 更细的 streaming、function/tool calling、多模态 part 类型,请以 Vertex AI 官方文档为准。