Integrations and generation

Integrations are the services you connect once and then use from inside the project: image generation, video, speech, transcription, and the web search an agent can reach for.

Chat models are not here — those are set up in Connect an LLM provider. A vendor you already use for chat appears in this catalog only for the things it does besides chat, and it arrives already connected.

The catalog

Open Integrations from the tray at the foot of the navigation rail. It is a catalog, not a form: search at the top, filters down the side, one card per service.

Cards are grouped by service, not by capability. A service that makes both images and video is one card with two tools inside it, because that is how you think about it — and because what you connect is the account, once.

Opening a card gives that service its own tab, so two services can sit side by side. Inside: its tools across the top, the form for the selected tool on the left, and that tool's runs on the right.

Keys

Each service says whether it is usable:

Two places hold keys, and it is worth knowing which is which. Note that two different surfaces are called Integrations: this catalog, opened from the nav rail, and the credentials section inside Settings.

Section Holds
Settings → LLM Providers Chat model keys, on each provider's own API Keys tab.
Settings → Integrations Keys for everything else — the services in this catalog.

Running a tool

Pick a tool, pick a model, write the prompt, and adjust whatever that model exposes: size, quality, count, duration, mode, a negative prompt. Only the knobs the chosen model actually has are shown.

Video tools that can start from a still image take a first frame — pick an existing file from the project. The service says what it accepts, and obvious mismatches are caught before sending rather than after the run — a picture Zero cannot make sense of is passed along rather than refused, since a wrong rejection is worse than a late failure.

Transcription

Not every tool starts from a prompt. Transcription — speech to text — starts from a file: instead of a prompt box you get a picker for an audio asset already in the project, and the run card shows the file's name where a prompt would have been.

The picker is narrowed to what the service accepts and states its size limit underneath. The output formats depend on the model, not only its quality: subtitles (.srt, .vtt) come from the classic Whisper model, and the newer ones return plain text or JSON. The result is saved into the project as an ordinary file, like any other generated artifact.

One model also offers translate, which returns English regardless of what was spoken. It is a switch on that model rather than a separate tool — same input, same output, only the language of the result differs. With it on, the language field means the language to produce rather than the one to expect, so anything other than English is refused with an explanation instead of being quietly ignored.

Runs

A run goes into the background. Its card shows the status and lets you stop it, restart it after a failure, or delete it. Finished results are saved into the project as ordinary files — nothing has to be downloaded from anywhere afterwards.

Deleting a run removes the run, not what it produced. Generated files are yours to keep or clear out.

Services Zero does not ship with

The catalog is not a fixed list. A service can be described by a file in the project — its endpoint, how it authenticates, what it takes and what it returns — and it then appears in the catalog alongside the built-in ones, marked by where it came from. Services shipped with Zero cannot be shadowed by a project file.