AI Lab
Experiment with AI like an engineer, not a chat session.
A developer workbench for designing, testing, comparing, and benchmarking AI calls across multiple providers, and composing them into multi-step workflows. It runs entirely on your own machine with your own API keys, and is open source under the MIT license so other developers can use it too.
Problem
Testing a prompt in a single provider's console until it seems to work isn't the same as knowing how it behaves — across models, across repeated runs, or as part of a larger multi-step workflow, that behavior stays invisible without tooling built to compare it.
Solution
AI Lab gives developers a local, provider-agnostic workbench for running the same request against OpenAI, Claude, Gemini, and Grok, comparing outputs and execution metrics side by side, benchmarking consistency across repeated runs, and composing individual requests into dependency-graph workflows — with provider credentials that stay local instead of passing through a hosted service.
Intended Users
Developers building AI-enabled applications who need to compare models and prompts, measure cost and latency tradeoffs, or design multi-step AI workflows, rather than testing changes by hand in a provider console.
Key Capabilities
- Runs entirely on your own machine, using your own API keys — credentials and data never pass through a hosted intermediary
- Side-by-side request execution across OpenAI, Claude, Gemini, and Grok
- Benchmarking for consistency, latency, token usage, and cost across repeated runs
- Multi-step AI workflows modeled as dependency graphs, with independent steps executing concurrently
- Full execution history, including resolved prompts, provider responses, and streaming timing
- Open source under the MIT license