You set the boundaries
Choose the environment, approved data, users, and model.
[SYSTEM: INITIALIZING]
Run private models on your hardware for less than public APIs cost.
Anything your team runs against a public AI service today can run against Epsilon instead: one endpoint, on your GPUs. You bring the model and the tools; we serve it, and nothing leaves the building.
The document assistant you already use, pointed at your own model. Contracts, client files and internal documents get searched, summarized and drafted from on your hardware, with nothing sent out.

You decide where the system runs, which model it uses, and what data it can access. Refactr deploys and maintains the serving layer so private AI costs less to run.
Choose the environment, approved data, users, and model.
In a private deployment, company work stays inside the environment you choose.
Deploy a compatible open model as-is, or fine-tune one for your workflow and approved data.
Start with the GPUs you have. Epsilon serves more work per card, so private AI costs less to run.
We deploy and maintain the serving layer: setup, monitoring, updates, and support. Not your models or apps.
How Refactr runs the stack
Standard serving leaves GPUs waiting between requests, and that idle time is what you pay for. Epsilon is the layer that keeps them working: the same cards produce more tokens for less GPU time, measured below. All of it runs inside your environment.

Same two RX 7900 XTX cards, same Qwen 3.8 (27B), measured under load. llama.cpp and vLLM ran bare on the cards; Epsilon ran on the same cards. Every figure below comes from those runs, so the only difference is the serving layer.
Epsilon needs fewer GPU-hours to produce the same output, so every million tokens costs less at any hourly rate. At $0.75 per GPU-hour that is $5.77 with Epsilon against $7.34 on vLLM and $8.59 on llama.cpp.
There are three other ways to get AI into the company: pay a public API per token, have your own team run a model, or rent managed inference in someone else’s cloud. Epsilon is the one where the model runs on your hardware and Refactr keeps it running.
Swipe to compare systems
How we start
Evaluate your models, throughput goals, and the GPUs you already control.
You set environment and access. We never touch models, documents, or code.
Install Epsilon on your GPUs to raise throughput and cut serving cost.
Check latency, throughput, and token cost. Expand, tune, or stop.