Inference × WorkDaddy: LLM Provider Integration (2026)
Use Inference as the engine behind your WorkDaddy agent team: 9 models available, context windows up to 125K tokens, connected in minutes with your own key.
Inference is one of 178 LLM providers supported by WorkDaddy’s model layer in 2026, in the Inference Clouds group. The catalog currently lists 9 models from Inference: 0 with explicit reasoning, 3 with vision input, and 8 with tool calling — the capability that matters most for agent work. The largest context window reaches 125K tokens.
At a glance
| Models available | 9 |
|---|---|
| Reasoning models | 0 |
| Vision models | 3 |
| Tool-calling models | 8 |
| Largest context window | 125K tokens |
| API endpoint | https://inference.net/v1 |
| API key variable | INFERENCE_API_KEY |
| SDK adapter | @ai-sdk/openai-compatible |
| Provider docs | inference.net ↗ |
Why run WorkDaddy agents on Inference in 2026
Tool calling, context size, and cost decide how well a provider drives an agent team. Inference brings 8 tool-calling models to WorkDaddy — enough to power the full loop of storefront edits, analytics queries, and content production — with 125K tokens of context at the top end for whole-codebase and long-report work. Because WorkDaddy is model-agnostic, you can route heavy reasoning to Inference’s strongest model and bulk work to its cheapest, inside one team.
Best Inference models for e-commerce agents (2026)
The current flagships from Inference in the WorkDaddy catalog, newest first:
Google Gemma 3 — Open Gemma instruction model for efficient chat and self-hosted deployments (125K context)
Qwen 2.5 7B Vision Instruct — Qwen vision-language model for visual reasoning, documents, and agent tasks (125K context)
Qwen 3 Embedding 4B — Embedding model for semantic search, retrieval, clustering, and ranking pipelines (32K context)
Llama 3.2 1B Instruct — Open Llama instruction model for multilingual chat, reasoning, and coding (16K context)
Llama 3.2 3B Instruct — Open Llama instruction model for multilingual chat, reasoning, and coding (16K context)
Llama 3.1 8B Instruct — Open Llama instruction model for multilingual chat, reasoning, and coding (16K context)
How to connect Inference to WorkDaddy
WorkDaddy talks to Inference at inference.net via the @ai-sdk/openai-compatible adapter. Add your credential (typically INFERENCE_API_KEY) in Settings → Models, pick a default model, and every agent can use it immediately — or add Inference as one engine among several and let tasks route to the best fit.
Featured models
| Model | Context | Reasoning | Released |
|---|---|---|---|
| Google Gemma 3 | 125K | No | 2025-01-01 |
| Qwen 2.5 7B Vision Instruct | 125K | No | 2025-01-01 |
| Qwen 3 Embedding 4B | 32K | No | 2025-01-01 |
| Llama 3.2 1B Instruct | 16K | No | 2025-01-01 |
| Llama 3.2 3B Instruct | 16K | No | 2025-01-01 |
| Llama 3.1 8B Instruct | 16K | No | 2025-01-01 |
How it works
Connect
In WorkDaddy, open Settings → Models, choose Inference, and paste your API key (INFERENCE_API_KEY). No key? Start on the managed gateway instead.
Put agents to work
Pick Google Gemma 3 or any of the 9 available models as your default — per-agent overrides let you match model to task.
Review and approve
Agent output stays draft-first regardless of the model: review diffs and approve actions exactly as before.
Frequently asked questions
Does WorkDaddy support Inference in 2026?
Yes — Inference is a supported LLM provider with 9 models in the WorkDaddy catalog, connected with your own API key via the @ai-sdk/openai-compatible adapter.
How many Inference models can I use with WorkDaddy?
9 models are listed for Inference, including 0 reasoning models and 3 vision-capable models. The flagship lineup above shows the newest.
What context window do Inference models offer?
Up to 125K tokens on the largest model — relevant for whole-repo storefront work and long analytics reports.
Is Inference the best LLM provider for e-commerce agents?
It depends on the task mix — Inference sits in the Inference Clouds group. WorkDaddy is model-agnostic, so the practical answer in 2026 is to combine providers: route each agent task to whichever model fits best.
Keep reading
NovitaAI
Inference Clouds
Nvidia
Inference Clouds
Venice AI
Inference Clouds
DigitalOcean
Inference Clouds
Explore more integration directories
Put the team to work on your store
Connect your storefront, analytics, and email stack, set a goal, and let the agents run the work end to end. Start on the free plan with your own model key.