Google Gemini
Auto-instrument the Google Gen AI SDK (google-genai on PyPI, @google/genai on npm). Every generateContent, generateContentStream and embedContent call inside an active run emits an llm_call event with model, token usage, cached and thinking tokens, latency, finish reason, and function calls. Cost is computed server-side from the rate card.
The older google-generativeai package and Vertex AI's own SDK are not covered. Calling Gemini through LangChain or LlamaIndex? Use the framework integration instead; it reports the same fields.
Install / enable
Python:
pip install "agentping-io[gemini]"
import agentping
from google import genai
agentping.init()
agentping.instrument_gemini()
client = genai.Client()
with agentping.run("support-triage", customer_id="acme-corp"):
reply = client.models.generate_content(
model="gemini-2.5-flash",
contents="classify this ticket",
)
instrument_gemini() patches google.genai.models.Models and AsyncModels, so client.models and client.aio.models are both covered. Idempotent and global; call it once at module load. Supported google-genai versions: 1.x and 2.x.
TypeScript:
npm install @agentping/sdk @google/genai
import { GoogleGenAI } from "@google/genai";
import * as agentping from "@agentping/sdk";
agentping.init({ apiKey: process.env.AGENTPING_API_KEY });
const run = agentping.run("answer-bot");
const gemini = agentping.instrumentGemini(
new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY }),
{ run },
);
const reply = await gemini.models.generateContent({
model: "gemini-2.5-flash",
contents: "hi",
});
await run.finish({ status: "success" });
instrumentGemini returns a wrapped client whose models namespace has the same return types as the original.
What's captured
| Field | Source |
|---|---|
provider |
gemini |
model |
response.model_version (the concrete version that served the call), falling back to the requested model |
input_tokens |
usage_metadata.prompt_token_count (already gross) |
output_tokens |
usage_metadata.candidates_token_count |
cached_input_tokens |
usage_metadata.cached_content_token_count |
reasoning_tokens |
usage_metadata.thoughts_token_count |
finish_reason |
candidates[0].finish_reason, as the enum name (STOP, MAX_TOKENS) |
tool_calls |
Number of function_call parts across the candidates |
latency_ms |
Wall-clock time of the call |
stream |
true on generateContentStream |
Gemini has no cache-creation count, so cache_creation_input_tokens is never set. Zero-valued fields are omitted.
embedContent emits an llm_call with kind: "embedding", the requested model and input_tokens from usage_metadata when the API returns it. A call that raises is recorded as an llm_call with status: "error", error and exception, then re-raised.
Streaming
Both SDKs capture generateContentStream (Python: generate_content_stream, sync and async). Usage arrives on the final chunk; the llm_call fires once the stream is fully consumed or closed, with stream: true.
with agentping.run("answer-bot"):
for chunk in client.models.generate_content_stream(
model="gemini-2.5-flash",
contents="explain briefly",
):
print(chunk.text, end="", flush=True)
# llm_call is emitted here
Source / notes
- Python:
agentping.instrument_gemini()in agent-ping-python - TypeScript:
instrumentGeminiin agent-ping-typescript
Default rates ship for Anthropic and OpenAI models only. Gemini calls land with their token counts and appear in the unpriced-models list until you add the model's per-million rates under Spend, Rate cards; from then on every new call is priced. See Spend.