Skip to content
Service

RAG systems & knowledge assistants

Most RAG demos fall over on real documents. We build the unglamorous parts properly: ingestion and chunking tuned to your content, hybrid retrieval instead of naive vector search, reranking, citation of the source passage, and an evaluation set so you can prove an answer got better rather than just different.

Scope this projectBook a 30-min call3–10 weeksPackages from $108

What you get

The outcomes we hold ourselves to on this work. If we cannot commit to them for your project, we say so during scoping rather than after the invoice.

  • Answers grounded in your documents, with the source passage cited
  • A measured retrieval quality baseline — not vibes
  • Hallucination handled explicitly: the system says when it does not know
  • Per-query cost and latency you can see and control

Deliverables

Everything on this list is handed over and documented. Anything outside it is quoted before it is built.

  • Ingestion pipeline with chunking and metadata strategy
  • Hybrid search: vector plus keyword, with reranking
  • Evaluation set and retrieval quality scoring
  • Answer synthesis with inline citations
  • Access control so users only retrieve what they may read
  • Cost, latency and token dashboards

Typical timeline

Most engagements of this type run 3–10 weeks end to end. Step four absorbs the difference between the short and long end of that range.

  1. Discovery call

    30 minutes, free

    We walk through what you are building, who it is for, and what has to be true for the project to count as a success. You leave with a rough scope, a rough number and an honest read on whether we are the right team.

    You receive: Written scope summary within 24 hours

  2. Proposal & fixed scope

    2–3 days

    A written proposal: milestones, deliverables, timeline, price and explicit exclusions. No hourly surprises — you approve a scope, and changes to it are quoted before any work starts.

    You receive: Proposal, contract and milestone schedule

  3. Architecture & design

    Week 1

    Data model, API contracts and infrastructure plan on paper before code. UI work starts in Figma so you approve screens rather than reviewing half-built pages.

    You receive: Architecture doc, ER diagram, approved designs

  4. Build in weekly sprints

    Bulk of the project

    Working software every week on a staging URL you can click through. A short written update each Friday: what shipped, what is next, anything blocking. You are never guessing where the project stands.

    You receive: Staging environment, weekly demo and written update

  5. Test, harden & launch

    Final 1–2 weeks

    Automated tests in CI, load and failure-path testing, security review, performance budgets, then a rehearsed deploy with a rollback plan. Launch day is uneventful by design.

    You receive: Test suite, CI pipeline, production deployment

  6. Handover & support

    Ongoing

    Documentation, a recorded walkthrough and a 30-day warranty on everything we shipped. If you want us to keep operating it, a maintenance retainer picks up from there.

    You receive: Docs, walkthrough video, 30-day warranty

Questions we get asked

Do we need to fine-tune a model?
Almost never at the start. Retrieval quality, chunking and prompt structure account for most of the gap between a bad assistant and a good one, and they are far cheaper to iterate on. We revisit fine-tuning only once retrieval is measured and still short.
Can it run on our own infrastructure?
Yes. Open-weight models on your own GPUs, or a hosted API — the retrieval layer is the same either way, so the decision is about compliance and cost, not architecture.

Need rag systems & knowledge assistants?

Tell us what you are working with — existing code, a blank repo, or a deadline you are worried about. Packages for this start at $108, and anything outside them gets a written scope and a number within 24 hours.

Replies within 4 business hours · No obligation · You keep the scope document