AI · SaaS
StudioX AI
Enterprise AI platform that grounds assistants in company knowledge and automates multi-step operational work.

Overview
Agents that execute real operational workflows against enterprise systems — with permission-aware retrieval, cost controls and full request tracing — deployable to cloud or on-premise.
The challenge
Enterprise teams had knowledge scattered across SharePoint, Confluence, Jira, ticketing systems and years of unstructured documents. Off-the-shelf assistants couldn't reach any of it — and the ones that could ignored the permission model, so any answer risked exposing documents the asker was never cleared to read. Meanwhile the operational work that actually consumed headcount was multi-step: read a mailbox, extract fields, query a knowledge base, generate a document, send it on. That needed automation, not conversation.
The solution
A platform built around three layers: a document ingestion pipeline that normalizes ~15 source types into a vector store, a retrieval layer that enforces per-user access control before the search runs, and a declarative workflow runtime where automations are authored as YAML and executed step by step. A single GraphQL API fronts roughly 100 feature domains on top of that foundation.
What we delivered
- Workflow runtime — YAML parsed into an executable step graph with ~26 step types: LLM processing, structured variable extraction, HTTP calls, knowledge-base queries, SQL extraction, conditionals, loops, scheduling, file generation (XLSX, PPTX, PDF) and email dispatch. Supports sub-workflow invocation with namespaced IDs, credential injection, variable interpolation and persisted execution state.
- Access-controlled retrieval — an LLM pass resolves the query against the tenant's semantic labels (aliased to opaque IDs before being shown to the model) and emits a metadata filter, so the vector search itself is scoped to what the user may read.
- Ingestion pipeline — a 10-stage pipeline (upload → linked-source sync → pre-processing → text conversion → extraction → embedding → retrieval prompt → deletion) dispatched across ~15 source types through a single strategy table, so adding a format means filling in cells rather than branching logic.
- Retrieval infrastructure — self-hosted Milvus behind a singleton client doing per-partition write locking and debounced flushes, with cross-encoder reranking (BGE, later mxbai) served from a Flask/PyTorch sidecar with GPU memory monitoring, batching and request queueing.
- LLM client — Azure OpenAI wrapper with token estimation (including image tokens), per-tenant budget enforcement before dispatch, bounded exponential-backoff retries, typed retryable vs. terminal errors, and credential refresh on auth failure.
- API layer — code-first GraphQL schema over ~100 feature domains, with a schema-builder tracing plugin capturing arguments, duration and errors for every root resolver, plus MongoDB driver instrumentation for slow-query detection.
- Background processing — ~90 recurring jobs plus an async task layer for long multi-step work (transcription, document comparison, policy generation), with pipeline-step chaining and retry on failure.
- Deployment — multi-stage Docker builds through GitLab CI to ECR and EKS, with separate staging and production accounts and per-customer image variants for on-premise installs.
Architecture
Clients → GraphQL (Pothos + Yoga) / REST webhooks / Socket.io
↓
Resolver layer (~100 feature domains)
↓
┌─────────┼──────────────┬──────────────────┐
↓ ↓ ↓ ↓
Workflow Retrieval Ingestion Job runner
runtime (ACL-aware) pipeline (Agenda)
↓ ↓ ↓ ↓
Azure Milvus + 10 stages × ~90 recurring
OpenAI reranker 15 source types + async tasks
↓
MongoDB · Redis · S3 / Azure BlobKey engineering decisions
- Declarative workflows over hardcoded ones. Automations are YAML documents parsed into a step graph, so new ones ship without a deploy. The runtime supports conditionals, loops, scheduling, sub-workflow invocation with namespaced step IDs, and human-in-the-loop prompts.
- Access control pushed into retrieval, not bolted on after. Filtering results post-search leaks existence; filtering pre-search doesn't. An LLM pass maps the question to the labels the asker is cleared for, and only those become the query filter.
- Self-hosted reranking. A cross-encoder sidecar reranks candidates from the vector search, keeping quality high without shipping document text to a third-party reranking API — which also made the on-premise deployments possible.
- Milvus over the managed vector DB. Migrated off a hosted vector service to self-hosted Milvus for cost control at volume and to support air-gapped installs.
- MongoDB-backed job coordination. Recurring jobs claim ownership through the shared job collection at boot, so horizontally scaled replicas don't double-process.
Tech stack
Engagement
2025 · AI Engineer
Results
A production platform serving multi-tenant SaaS customers and on-premise enterprise installations from one codebase, where operations teams author their own automations in YAML and assistants answer from company knowledge without crossing permission boundaries.
Lessons learned
- Retrieval quality is mostly a data-engineering problem. The wins came from normalizing messy enterprise sources properly, reranking hard, and scoping the search by permission before it ran — not from prompt tuning. And making access control part of the query rather than a filter on the results is what turned a demo into something a regulated customer would actually deploy.
Have a hard problem like this?
Let's talk about what you're building.