What we build
- Fine-tuned models: LoRA and QLoRA training on open-weight models for your domain, with an evaluation harness that compares every candidate against the base model on your own benchmark.
- Retrieval systems: self-hosted RAG over your documents with citation-grounded answers, document-level access control, and a full audit log. Nothing leaves your perimeter.
- Document intelligence: OCR and extraction pipelines that turn invoices, contracts, and customs documents into validated structured data, with human review where confidence is low.
- Inference optimization: quantization, distillation, caching, and serving with vLLM, so the model you ship is one you can afford to run.
- Supervised agents for back-office queues: state-machine agents with approval steps, action logs, and replay. Automation rates you can measure, not demos.
How we work
Every AI engagement starts with the evaluation set, not the model. We define what good output looks like with your team, build the benchmark, and only then train. If fine-tuning does not beat the base model plus retrieval on your data, we tell you and you keep the benchmark.
Optimization is a tradeoff we make explicit: speed against quality against memory. You see the numbers for each configuration before anything ships.
Models run where your data lives: your cloud account, your cluster, or our EU hosting platform. We do not route your data through third-party APIs unless you choose that deliberately.
Compliance is part of the build
For systems the EU AI Act treats as high risk, we produce the technical documentation, logging, and monitoring evidence alongside the system itself. Your legal team gets a paper trail, not a promise.