AI Technologies and Techniques
The difference between an AI demo and a production system is technique. Software Depo builds with the methods that make AI reliable in real business operations — retrieval grounding, disciplined agent architectures, rigorous evaluation, and security at the boundary between AI and your systems.
Below are the core techniques we engineer with. Each one exists to make AI accurate, safe, and maintainable — not just impressive in a first demo.
Core Techniques
Production AI Engineering
Software Depo engineers AI with retrieval grounding, agent architectures, the Model Context Protocol, evaluation, and observability — the techniques that make AI reliable, citable, and safe in production. Read our AI development insights or start a conversation.
Which Technique Fixes Which Failure?
Start from the symptom, not the technique. State the symptom precisely and the choice narrows on its own.
- Confidently wrong about your own data: retrieval grounding. Check whether the retrieval step returned the right passage before blaming the model.
- The right document comes back and the answer still misses: chunking, metadata, and a reranking pass.
- The agent calls the wrong tool or invents arguments: tool schemas and descriptions, covered under Tool and Function Calling.
- Fine in the demo, drifting in production: evaluation sets and observability.
- Fine until someone feeds it hostile text: permission scoping and guardrails.
Retrieval Is a Search Problem First
Pure vector similarity is weak on exact identifiers: part numbers, ticket references, contract clauses, people’s names. If your corpus is full of those, a larger model does not address the failure. Hybrid keyword and vector search, metadata filters that respect who is allowed to see what, and a reranking pass act on the step that is losing the document.
Citations are part of the technique. They do not make an answer correct; they make it checkable, and an answer nobody can check has its errors surface somewhere downstream instead. Retrieval-Augmented Generation and Embeddings and Semantic Search cover how the index gets built. RAG vs Fine-Tuning covers the choice before anyone proposes training on your documents.
Prompt Injection as an Engineering Constraint
Any system that reads text from outside and can act on internal systems carries this exposure, and no amount of wording in a system prompt closes it. The defenses are architectural: a separate least-privilege credential for each tool, a clear boundary between retrieved content and instructions, human approval before writes and irreversible actions, allowlists on destinations, and logging that lets you trace a chain backwards after the fact.
Testing narrows the surface. It should not be mistaken for proof that the surface is gone.
Evaluation Makes Everything Else Measurable
Without a scored set of cases, every change is a guess and every improvement is anecdote. The set matters less than the habit of running it before a change goes out.
The other half is production. Log the input, the retrieved context, the tool calls, and the output together, so a complaint from a user becomes a trace someone can read instead of a story someone has to reconstruct. Evaluation and Testing and Observability and Monitoring cover the two halves separately.
How Much of This You Actually Need
Not all of it. A read-only assistant over a small, stable document set needs solid retrieval and very little else, and adding governance machinery to it mostly adds friction. An agent with write access to customer records needs the full set: scoped permissions, approval steps, evaluation, and an audit trail. Match the machinery to the blast radius, then revisit the decision when the blast radius grows.