All projects & case studies
AISaaSAutomation

AI · SaaS

StudioX AI

Enterprise AI platform that grounds assistants in company knowledge and automates multi-step operational work.

StudioX AI
Met client companies directly, proposed AI solutions for their workflows, and took them from discovery to production — legal document work for law firms, data analysis for IoT companies, document chatbots for employees and customers, and sales analytics
Built a declarative workflow runtime that lets teams automate multi-step operational work — reading mailboxes, extracting fields, querying knowledge, generating documents and sending them — without shipping code for each new automation
Designed a permission-aware retrieval layer so assistants answer only from documents the asker is cleared to see, which is what made the platform viable for regulated customers

Overview

Agents that execute real operational workflows against enterprise systems — with permission-aware retrieval, cost controls and full request tracing — deployable to cloud or on-premise.

The challenge

Enterprise teams had knowledge scattered across SharePoint, Confluence, Jira, ticketing systems and years of unstructured documents. Off-the-shelf assistants couldn't reach any of it — and the ones that could ignored the permission model, so any answer risked exposing documents the asker was never cleared to read. Meanwhile the operational work that actually consumed headcount was multi-step: read a mailbox, extract fields, query a knowledge base, generate a document, send it on. That needed automation, not conversation.

The solution

A platform built around three layers: a document ingestion pipeline that normalizes ~15 source types into a vector store, a retrieval layer that enforces per-user access control before the search runs, and a declarative workflow runtime where automations are authored as YAML and executed step by step. A single GraphQL API fronts roughly 100 feature domains on top of that foundation.

What we delivered

  • Workflow runtime — YAML parsed into an executable step graph with ~26 step types: LLM processing, structured variable extraction, HTTP calls, knowledge-base queries, SQL extraction, conditionals, loops, scheduling, file generation (XLSX, PPTX, PDF) and email dispatch. Supports sub-workflow invocation with namespaced IDs, credential injection, variable interpolation and persisted execution state.
  • Access-controlled retrieval — an LLM pass resolves the query against the tenant's semantic labels (aliased to opaque IDs before being shown to the model) and emits a metadata filter, so the vector search itself is scoped to what the user may read.
  • Ingestion pipeline — a 10-stage pipeline (upload → linked-source sync → pre-processing → text conversion → extraction → embedding → retrieval prompt → deletion) dispatched across ~15 source types through a single strategy table, so adding a format means filling in cells rather than branching logic.
  • Retrieval infrastructure — self-hosted Milvus behind a singleton client doing per-partition write locking and debounced flushes, with cross-encoder reranking (BGE, later mxbai) served from a Flask/PyTorch sidecar with GPU memory monitoring, batching and request queueing.
  • LLM client — Azure OpenAI wrapper with token estimation (including image tokens), per-tenant budget enforcement before dispatch, bounded exponential-backoff retries, typed retryable vs. terminal errors, and credential refresh on auth failure.
  • API layer — code-first GraphQL schema over ~100 feature domains, with a schema-builder tracing plugin capturing arguments, duration and errors for every root resolver, plus MongoDB driver instrumentation for slow-query detection.
  • Background processing — ~90 recurring jobs plus an async task layer for long multi-step work (transcription, document comparison, policy generation), with pipeline-step chaining and retry on failure.
  • Deployment — multi-stage Docker builds through GitLab CI to ECR and EKS, with separate staging and production accounts and per-customer image variants for on-premise installs.

Architecture

Clients → GraphQL (Pothos + Yoga) / REST webhooks / Socket.io
             ↓
        Resolver layer (~100 feature domains)
             ↓
   ┌─────────┼──────────────┬──────────────────┐
   ↓         ↓              ↓                  ↓
Workflow  Retrieval     Ingestion          Job runner
runtime   (ACL-aware)   pipeline           (Agenda)
   ↓         ↓              ↓                  ↓
Azure    Milvus +      10 stages ×        ~90 recurring
OpenAI   reranker      15 source types     + async tasks
             ↓
     MongoDB · Redis · S3 / Azure Blob

Key engineering decisions

  • Declarative workflows over hardcoded ones. Automations are YAML documents parsed into a step graph, so new ones ship without a deploy. The runtime supports conditionals, loops, scheduling, sub-workflow invocation with namespaced step IDs, and human-in-the-loop prompts.
  • Access control pushed into retrieval, not bolted on after. Filtering results post-search leaks existence; filtering pre-search doesn't. An LLM pass maps the question to the labels the asker is cleared for, and only those become the query filter.
  • Self-hosted reranking. A cross-encoder sidecar reranks candidates from the vector search, keeping quality high without shipping document text to a third-party reranking API — which also made the on-premise deployments possible.
  • Milvus over the managed vector DB. Migrated off a hosted vector service to self-hosted Milvus for cost control at volume and to support air-gapped installs.
  • MongoDB-backed job coordination. Recurring jobs claim ownership through the shared job collection at boot, so horizontally scaled replicas don't double-process.

Tech stack

Azure OpenAI / LLMsAI AgentsRAGMilvusGraphQLTypeScriptPythonKubernetes

Engagement

2025 · AI Engineer

Results

A production platform serving multi-tenant SaaS customers and on-premise enterprise installations from one codebase, where operations teams author their own automations in YAML and assistants answer from company knowledge without crossing permission boundaries.

Lessons learned

  • Retrieval quality is mostly a data-engineering problem. The wins came from normalizing messy enterprise sources properly, reranking hard, and scoping the search by permission before it ran — not from prompt tuning. And making access control part of the query rather than a filter on the results is what turned a demo into something a regulated customer would actually deploy.

Have a hard problem like this?

Let's talk about what you're building.