Cohere

Auto-instrument the Cohere v2 client (cohere on PyPI, cohere-ai on npm). Every chat, chat_stream and embed call inside an active run emits an llm_call event with model, billed token counts, latency, finish reason, and tool calls. Cost is computed server-side from the rate card.

Only the v2 client (ClientV2 / CohereClientV2) is instrumented. The legacy v1 Client is not touched.

Install / enable

Python:

pip install "agentping-io[cohere]"
import agentping
import cohere

agentping.init()
agentping.instrument_cohere()

client = cohere.ClientV2()

with agentping.run("support-triage", customer_id="acme-corp"):
    reply = client.chat(
        model="command-a-03-2025",
        messages=[{"role": "user", "content": "classify this ticket"}],
    )

instrument_cohere() patches cohere.ClientV2 and cohere.AsyncClientV2. Idempotent and global; call it once at module load. Supported cohere versions: 5.x.

TypeScript:

npm install @agentping/sdk cohere-ai
import { CohereClientV2 } from "cohere-ai";
import * as agentping from "@agentping/sdk";

agentping.init({ apiKey: process.env.AGENTPING_API_KEY });

const run = agentping.run("answer-bot");
const cohere = agentping.instrumentCohere(
  new CohereClientV2({ token: process.env.COHERE_API_KEY }),
  { run },
);

const reply = await cohere.chat({
  model: "command-a-03-2025",
  messages: [{ role: "user", content: "hi" }],
});

await run.finish({ status: "success" });

instrumentCohere returns a wrapped client covering chat, chatStream and embed.

What's captured

Field Source
provider cohere
model The requested model. Cohere's response does not echo it, so the request is the only source.
input_tokens usage.billed_units.input_tokens
output_tokens usage.billed_units.output_tokens
finish_reason finish_reason
tool_calls Number of message.tool_calls
latency_ms Wall-clock time of the call
stream true on chat_stream

Token counts are the billed units, which is what Cohere charges for, rather than the raw tokens block. Cohere reports no cache or reasoning breakdown, so those fields are never set.

embed emits an llm_call with kind: "embedding", the requested model and input_tokens from meta.billed_units. A call that raises is recorded as an llm_call with status: "error", error and exception, then re-raised.

Streaming

Both SDKs capture chat_stream (chatStream in TypeScript). Usage and finish reason arrive on the terminal message-end event; each tool-call-start event counts one tool call. One llm_call with stream: true fires once the stream is fully consumed or closed.

with agentping.run("answer-bot"):
    for event in client.chat_stream(
        model="command-a-03-2025",
        messages=[{"role": "user", "content": "explain briefly"}],
    ):
        if event.type == "content-delta":
            print(event.delta.message.content.text, end="", flush=True)
    # llm_call is emitted here

Source / notes

Default rates ship for Anthropic and OpenAI models only. Cohere calls land with their billed token counts and appear in the unpriced-models list until you add the model's per-million rates under Spend, Rate cards; from then on every new call is priced. See Spend.