> ## Documentation Index
> Fetch the complete documentation index at: https://braintrust.dev/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# API reference

> Functions and configuration options found in the Braintrust Go SDK.

export const feature_0 = "Go remote evals"

export const verb_0 = "are"

This page covers the key APIs in the Braintrust Go SDK. For setup, see the [Quickstart](/docs/sdks/go/quickstart). For the complete reference, see [pkg.go.dev](https://pkg.go.dev/github.com/braintrustdata/braintrust-sdk-go).

## Tracing

Tracing records what your application does as spans you can inspect in Braintrust. The recommended way to capture AI calls is auto-instrumentation: use the `trace/contrib` packages to instrument supported provider libraries, either at build time with Orchestrion or with runtime middleware (see [Go SDK integrations](/docs/sdks/go/sdk-integrations)). Tracing is built on OpenTelemetry, so you trace your own code with the standard OpenTelemetry API. The APIs below create the client, trace your own code, and link to your traces.

### `braintrust.New`

Creates a Braintrust client and configures the OpenTelemetry pipeline that exports spans to Braintrust. Call it once on startup, passing your `TracerProvider` and any [options](#configuration).

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import (
	"github.com/braintrustdata/braintrust-sdk-go"
	"go.opentelemetry.io/otel"
	"go.opentelemetry.io/otel/sdk/trace"
)

tp := trace.NewTracerProvider()
otel.SetTracerProvider(tp)

client, err := braintrust.New(tp, braintrust.WithProject("My project"))
if err != nil {
	log.Fatal(err)
}
```

Returns: `(*braintrust.Client, error)`.

`braintrust.New` reads `BRAINTRUST_API_KEY` from the environment. Configure the rest with functional options or environment variables (see [Configuration](#configuration)).

Because tracing is built on OpenTelemetry, you trace your own application code with the standard OpenTelemetry API, and traced AI calls nest under your spans.

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
ctx, span := otel.Tracer("my-app").Start(ctx, "process-request")
defer span.End()
```

### `Client.Permalink`

Builds a Braintrust UI URL for a span, so you can link straight to a trace from your own logs or app.

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
url := client.Permalink(span)
```

Returns: `string`.

## Evaluations

An evaluation runs your task over a set of cases, scores each output, and logs the results to an experiment, which is how you measure quality and catch regressions as you change prompts or models.

The recommended pattern is to define an eval once with `braintrust.NewEval`, then call `Run` with any dataset. The same definition works for local runs, `bt eval <dir>` runs from the command line, and [remote eval](/docs/evaluate/remote-evals) runs triggered from the playground.

### `braintrust.NewEval`

Creates a runnable `*eval.Eval` by combining a client with an eval definition. Call `Run` on it to execute the evaluation.

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import (
	"context"

	"github.com/braintrustdata/braintrust-sdk-go"
	"github.com/braintrustdata/braintrust-sdk-go/eval"
)

e := braintrust.NewEval(client, &eval.Eval[string, string]{
	Name: "classify",
	Task: eval.T(func(ctx context.Context, input string) (string, error) {
		return classify(input), nil
	}),
	Scorers: []eval.Scorer[string, string]{exactMatch},
	ProjectName: "my-project",
})

_, err := e.Run(ctx, eval.RunOpts[string, string]{
	Dataset: eval.NewDataset([]eval.Case[string, string]{
		{Input: "apple", Expected: "fruit"},
	}),
})
```

Returns: `*eval.Eval[I, R]`. `Run` returns `(*eval.Result, error)`.

`eval.Eval[I, R]` fields:

* **`Name`** (`string`, required): eval name, used as the default experiment name and as the registration key for remote evals.
* **`Task`** (`eval.TaskFunc[I, R]`, required): the function under evaluation. Wrap a plain `func(ctx, input) (output, error)` with `eval.T`, or use `eval.TaskWithHooks` if the task needs to read parameters or other run metadata.
* **`Scorers`** (`[]eval.Scorer[I, R]`): scoring functions applied to each case. Provide `Scorers`, `Classifiers`, or both.
* **`Classifiers`** (`[]eval.Classifier[I, R]`): classifiers applied to each case.
* **`ParameterSchema`** (`eval.ParameterSchema`): declares parameters the eval accepts. Each entry becomes a configurable control in the Braintrust playground (see [Parameters](#parameters)).
* **`Dataset`** (`eval.Dataset[I, R]`): default dataset for the eval. A playground run supplies its own dataset and overrides this. `bt eval <dir>` and a direct `Run` call use it when no dataset is supplied at call time.
* **`ProjectName`** (`string`): Braintrust project for this eval. Falls back to the client's configured project.

`eval.RunOpts[I, R]` fields (passed to `Run`):

* **`Dataset`** (`eval.Dataset[I, R]`): test cases for this run. Defaults to `Eval.Dataset`; required if the eval defines none.
* **`Experiment`** (`string`): experiment name. Defaults to `Eval.Name`.
* **`ProjectName`** (`string`): overrides the project for this run.
* **`ProjectID`** (`string`): identifies the project by ID instead of by name. Takes precedence over `ProjectName` when set.
* **`Tags`** (`[]string`): tags to apply to the experiment.
* **`Metadata`** (`eval.Metadata`): metadata to attach to the experiment.
* **`Update`** (`bool`): append to an existing experiment with the same name. Defaults to `false`.
* **`Parallelism`** (`int`): number of goroutines. Defaults to `1`.
* **`Quiet`** (`bool`): suppress result output. Defaults to `false`.

### `braintrust.NewEvaluator`

Creates an evaluator for input type `I` and result type `R`, bound to a client. Call `Run` on it with the cases, task, and scorers to execute the evaluation and log an experiment.

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import (
	"context"

	"github.com/braintrustdata/braintrust-sdk-go"
	"github.com/braintrustdata/braintrust-sdk-go/eval"
)

evaluator := braintrust.NewEvaluator[string, string](client)

_, err := evaluator.Run(context.Background(), eval.Opts[string, string]{
	Experiment: "answers-v1",
	Dataset: eval.NewDataset([]eval.Case[string, string]{
		{Input: "How do I reset my password?", Expected: "Use the account recovery flow."},
		{Input: "How do I export my data?", Expected: "Open Settings and choose Export."},
	}),
	Task: eval.T(answerQuestion),
	Scorers: []eval.Scorer[string, string]{
		eval.NewScorer("exact_match", func(_ context.Context, r eval.TaskResult[string, string]) (eval.Scores, error) {
			v := 0.0
			if r.Output == r.Expected {
				v = 1.0
			}
			return eval.S(v), nil
		}),
	},
})
```

Returns: `*eval.Evaluator[I, R]`. `Run` returns `(*eval.Result, error)`.

`eval.Opts[I, R]` fields:

* **`Experiment`** (`string`, required): experiment name.
* **`Dataset`** (`eval.Dataset[I, R]`, required): the cases to run. Build an in-memory one with `eval.NewDataset` (see [Datasets](#datasets)).
* **`Task`** (`eval.TaskFunc[I, R]`, required): the function under test. Wrap a plain `func(ctx, input) (output, error)` with `eval.T`.
* **`Scorers`** (`[]eval.Scorer[I, R]`): scorers to apply to each case. Provide `Scorers`, `Classifiers`, or both.
* **`Classifiers`** (`[]eval.Classifier[I, R]`): classifiers to apply to each case. Provide `Scorers`, `Classifiers`, or both.
* **`ProjectName`** (`string`): project to log to. Defaults to the client's configured project.
* **`ProjectID`** (`string`): project ID. When set, takes precedence over `ProjectName` and uses the project as-is rather than creating it.
* **`Tags`** (`[]string`): tags to apply to the experiment.
* **`Metadata`** (`eval.Metadata`): metadata to attach to the experiment.
* **`Update`** (`bool`): append to an existing experiment with the same name. Defaults to `false`.
* **`Parallelism`** (`int`): number of goroutines. Defaults to `1`.
* **`TrialCount`** (`int`): number of times to run each case. Defaults to `1`.
* **`Quiet`** (`bool`): suppress result output. Defaults to `false`.

### `eval.NewScorer`

Creates a scorer from a function. A scorer measures how good the task's output is, returning one or more named scores per case.

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
scorer := eval.NewScorer("exact_match", func(_ context.Context, r eval.TaskResult[string, string]) (eval.Scores, error) {
	if r.Output == r.Expected {
		return eval.S(1.0), nil
	}
	return eval.S(0.0), nil
})
```

Returns: `eval.Scorer[I, R]`. The score function receives an `eval.TaskResult[I, R]` (with `Input`, `Output`, `Expected`, and `Metadata`) and returns `eval.Scores`. Use `eval.S` to build a single score.

### `eval.NewClassifier`

Creates a classifier from a function. Use a classifier to categorize output instead of scoring it numerically.

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
classifier := eval.NewClassifier("topic", func(_ context.Context, r eval.TaskResult[string, string]) (eval.Classifications, error) {
	return eval.Classifications{{ID: "billing", Label: "Billing"}}, nil
})
```

Returns: `eval.Classifier[I, R]`.

### `eval.TaskWithHooks`

Creates a task function that receives `*eval.TaskHooks`, which gives access to parameters, metadata, tags, and the current spans. Use it when your task needs to read parameter values that were configured in the playground or passed via `RunOpts.Parameters`.

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
task := eval.TaskWithHooks(func(ctx context.Context, input string, hooks *eval.TaskHooks) (string, error) {
	model := hooks.Parameters.String("model")
	return classify(input, model), nil
})
```

Returns: `eval.TaskFunc[I, R]`. The hooks give access to `Parameters`, `Metadata`, `Tags`, `TrialIndex`, `TaskSpan`, and `EvalSpan`. Use `eval.T` instead when the task doesn't need hooks.

## Parameters

Parameters let you declare configurable options on an eval. When the eval runs from the Braintrust playground via a [remote eval](/docs/evaluate/remote-evals), each declared parameter becomes a control in the UI. For a local run, the task receives the declared defaults.

Declare parameters with `eval.ParameterSchema` on the `eval.Eval` definition:

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
e := braintrust.NewEval(client, &eval.Eval[string, string]{
	Name: "classify",
	ParameterSchema: eval.ParameterSchema{
		"model": {
			Type:        eval.ParameterTypeModel,
			Default:     "gpt-5-mini",
			Description: "Model to use for classification",
		},
		"threshold": {
			Type:    "number",
			Default: 0.5,
		},
	},
	Task: eval.TaskWithHooks(func(ctx context.Context, input string, hooks *eval.TaskHooks) (string, error) {
		model := hooks.Parameters.String("model")
		threshold := hooks.Parameters.Float64("threshold")
		return classify(input, model, threshold), nil
	}),
	// ...
})
```

`eval.ParameterSchema` is `map[string]eval.ParameterDef`. Each `eval.ParameterDef` has:

* **`Type`** (`string`): `"model"` renders a model picker. `"prompt"` renders a prompt editor. `"string"`, `"number"`, `"integer"`, `"boolean"` render plain inputs. Use the `eval.ParameterTypeModel` and `eval.ParameterTypePrompt` constants for the two dedicated controls.
* **`Default`** (`any`): used when no value is supplied for this run. This is also what a local run sees.
* **`Description`** (`string`): shown alongside the control in the playground.

`eval.Parameters` (the resolved values, available as `hooks.Parameters`) has typed accessors that never panic on a type mismatch:

* **`Parameters.String(name)`** → `string`
* **`Parameters.Int(name)`** → `int`
* **`Parameters.Float64(name)`** → `float64`
* **`Parameters.Bool(name)`** → `bool`
* **`Parameters.Get(name)`** → `(any, bool)`
* **`Parameters.Has(name)`** → `bool`

## Remote evals

Remote evals let the Braintrust playground trigger your Go eval code on your own infrastructure. The `evalrunner` package turns a Go binary into a target that the [`bt` CLI](/docs/reference/cli/eval) can drive.

<Warning>
  {feature_0} {verb_0} in [public preview](/docs/feature-lifecycle) and can change before reaching general availability.
</Warning>

### `evalrunner.New` and `evalrunner.RegisterEval`

Create a runner and register your evals. The runner reads the `bt` environment variables, dispatches the right eval, and streams results back.

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
package main

import (
	"context"

	"github.com/braintrustdata/braintrust-sdk-go/eval"
	"github.com/braintrustdata/braintrust-sdk-go/evalrunner"
)

func main() {
	r := evalrunner.New()

	evalrunner.RegisterEval(r, &eval.Eval[string, string]{
		Name: "classify",
		Task: eval.TaskWithHooks(func(ctx context.Context, input string, hooks *eval.TaskHooks) (string, error) {
			model := hooks.Parameters.String("model")
			return classify(input, model), nil
		}),
		Scorers: []eval.Scorer[string, string]{exactMatch},
		ParameterSchema: eval.ParameterSchema{
			"model": {Type: eval.ParameterTypeModel, Default: "gpt-5-mini"},
		},
		ProjectName: "my-project",
	})

	evalrunner.Main(r)
}
```

Then run from the command line:

```bash theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
# Run all evals locally
bt eval ./cmd/evals

# Start the dev server so the playground can trigger runs
bt eval --dev --language go ./cmd/evals
```

`evalrunner.New` accepts `evalrunner.Option` values:

* **`evalrunner.WithLogger(l logger.Logger)`**: custom logger. Logs go to stderr. Stdout is reserved for the `bt` protocol.
* **`evalrunner.WithTracerProvider(tp *sdktrace.TracerProvider)`**: shared OpenTelemetry provider. Supply one when instrumented code (LLM clients, custom spans) should appear in the same trace as eval spans. When nil, a per-run provider is created and shut down on exit.

`evalrunner.RegisterEval[I, R any](r *Runner, ev *eval.Eval[I, R])` registers an eval by its `Name`. Registering two evals under the same name replaces the first. `evalrunner.Main(r *Runner)` dispatches and exits. Use `evalrunner.Run(ctx, r)` instead if you need to handle the error yourself.

`r.Mode()` returns the current `evalrunner.Mode`: `ModeList`, `ModeEval`, `ModeBatch`, or `ModeInspect`. Use it to skip expensive setup when `bt` only wants the list of registered evals:

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
if r.Mode() != evalrunner.ModeList {
	warmCaches()
}
```

## Datasets

A dataset is the set of cases an evaluation runs against. Define cases inline in memory, or manage datasets in Braintrust through the [API client](#api-client).

### `eval.NewDataset`

Groups cases into an in-memory dataset you pass to `Evaluator.Run`, as an alternative to loading one from Braintrust.

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
dataset := eval.NewDataset([]eval.Case[string, string]{
	{Input: "How do I reset my password?", Expected: "Use the account recovery flow."},
	{Input: "How do I export my data?", Expected: "Open Settings and choose Export."},
})
```

Returns: `eval.Dataset[I, R]`.

Each `eval.Case[I, R]` has an `Input` and optional `Expected`, `Tags`, `Metadata`, and `TrialCount`.

## Attachments

When your traces involve binary content like images or PDFs, log it as an attachment so it appears in Braintrust instead of as an opaque blob. When you trace AI calls, Braintrust automatically converts base64 attachments in provider messages into uploaded attachments, so you rarely need the APIs below for instrumented calls. Reach for them when you're attaching binary content to a span yourself.

### `attachment.From*`

Creates an attachment from bytes, a file, or a URL.

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import "github.com/braintrustdata/braintrust-sdk-go/trace/attachment"

att, err := attachment.FromFile("image/png", "chart.png")
```

Constructors:

* **`attachment.FromBytes(contentType string, data []byte)`** → `*attachment.Attachment`: from raw bytes.
* **`attachment.FromFile(contentType string, path string)`** → `(*attachment.Attachment, error)`: reads a file.
* **`attachment.FromURL(url string)`** → `(*attachment.Attachment, error)`: fetches a URL and uses the response content type.
* **`attachment.FromReader(contentType string, r io.Reader)`** → `*attachment.Attachment`: from an `io.Reader`.

## API client

For direct access to the Braintrust REST API, use the `api` package. Reach for it to manage projects, experiments, datasets, and functions programmatically, beyond what the higher-level APIs above cover.

### `api.NewClient`

Creates a REST API client from an API key.

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
import "github.com/braintrustdata/braintrust-sdk-go/api"

client := api.NewClient(os.Getenv("BRAINTRUST_API_KEY"))
```

Returns: `*api.API`.

Namespaces:

* **`client.Projects()`**: project management.
* **`client.Experiments()`**: experiment management.
* **`client.Datasets()`**: dataset management. Methods include `Create`, `Insert`, `InsertEvents`, `Delete`, `Fetch`, and `Query`.
* **`client.Functions()`**: function management, including `Invoke(ctx, functionID, input)` to call a deployed function.

## Configuration

Configure the client with functional options passed to `braintrust.New`, or with environment variables.

```go #skip-compile theme={"theme":{"light":"github-light","dark":"github-dark-dimmed"}}
client, err := braintrust.New(tp,
	braintrust.WithProject("My project"),
	braintrust.WithBlockingLogin(true),
)
```

Client options:

* **`braintrust.WithAPIKey(apiKey string)`**: API key. Defaults to `BRAINTRUST_API_KEY`.
* **`braintrust.WithAPIURL(apiURL string)`**: Braintrust API URL.
* **`braintrust.WithAppURL(appURL string)`**: Braintrust app URL, used for permalinks.
* **`braintrust.WithOrgName(orgName string)`**: organization name, useful when credentials can access multiple orgs.
* **`braintrust.WithProject(projectName string)`**: project that receives exported spans.
* **`braintrust.WithProjectID(projectID string)`**: project ID. Takes precedence over the project name.
* **`braintrust.WithBlockingLogin(enabled bool)`**: log in synchronously during `New` instead of in the background.
* **`braintrust.WithExporter(exporter trace.SpanExporter)`**: supply a custom span exporter. Intended for testing.
* **`braintrust.WithEnableTraceConsoleLog(enabled bool)`**: print spans to the console.
* **`braintrust.WithFilterAISpans(enabled bool)`**: export only AI-related spans.
* **`braintrust.WithEnvironment(environmentType string, name ...string)`**: set span-origin environment provenance. Overrides auto-detection from CI and server environment variables.

### Environment variables

* **`BRAINTRUST_API_KEY`** (required): Braintrust API key.
* **`BRAINTRUST_API_URL`**: Braintrust API URL. Defaults to `https://api.braintrust.dev`.
* **`BRAINTRUST_APP_URL`**: Braintrust app URL, used for permalinks. Defaults to `https://www.braintrust.dev`.
* **`BRAINTRUST_DEFAULT_PROJECT`**: project that traced spans route to. Defaults to `default-go-project`.
* **`BRAINTRUST_DEFAULT_PROJECT_ID`**: project UUID. Takes precedence over the project name.
* **`BRAINTRUST_ORG_NAME`**: organization name, useful when credentials can access multiple orgs.
* **`BRAINTRUST_OTEL_FILTER_AI_SPANS`**: set to `true` to export only AI-related spans.
* **`BRAINTRUST_ENABLE_TRACE_CONSOLE_LOG`**: set to `true` to print spans to the console.
* **`BRAINTRUST_OTEL_ENABLE_BUILTIN_ADK_TRACES`**: set to `true` to export spans from Google ADK's built-in telemetry. Defaults to `false`.
* **`BRAINTRUST_AUTO_CONVERT_AI_ATTACHMENTS`**: scan spans for base64 attachments and upload them to object storage. Defaults to `true`; set to `false` to disable.
* **`BRAINTRUST_DEBUG`**: set to `true` to enable SDK debug logging. Defaults to `false`.
* **`BRAINTRUST_BLOCKING_LOGIN`**: set to `true` to log in synchronously at startup.
* **`BRAINTRUST_ENVIRONMENT_TYPE`**: environment type for span-origin provenance (for example, `ci` or `server`). Detected automatically from common CI and server runtimes when unset.
* **`BRAINTRUST_ENVIRONMENT_NAME`**: environment name for span-origin provenance (for example, `github_actions` or `aws_lambda`). Detected automatically when unset.
