Author

AI Mastery Architect

Specializing in RAG infrastructure, GPU optimization, and agentic workflows. Leading the technical research and engineering tools development at AI Mastery.

AMA
AI Mastery ArchitectLead Systems Engineer

Specializing in RAG infrastructure, GPU optimization, and agentic workflows. Leading the technical research and engineering tools development at AI Mastery.

RAGCUDALLM OpsAgentic Systems

Recent work

Article
architect • 2026-05-25T09:00:00Z
AgentsRAGInformation RetrievalAI ResearchAgentic AI

Why AI Agents Need a Terminal, Not Just a Vector Database

Researchers propose Direct Corpus Interaction (DCI) — a technique that gives AI agents raw bash-like terminal access to search corpora, outperforming dense retrievers by over 30 points on multi-hop QA while cutting costs by 29%.

Read more →
guides
architect • 2026-05-25T09:00:00Z
Local LLMsOllamallama.cppRAGDockerGGUFLLM Engineering

The Complete Developer Guide to Running LLMs Locally: From Ollama to Production

Everything you need to run LLMs on your own hardware in 2026: VRAM sizing, model formats, an 8-tool comparison table, a full local RAG pipeline, and Docker production deployment with GPU passthrough and Nginx auth.

Read more →
guides
architect • 2026-05-25
AgentsEvent-Driven ArchitectureMulti-Agent SystemsEnterprise AIAI Infrastructure

Event-Driven Architecture for Agentic AI: The Architect's Guide

A comprehensive architectural guide to designing resilient, real-time agentic AI systems using event-driven architecture — covering loose coupling, fault isolation, reference architecture, and governance patterns.

Read more →
guides
architect • 2026-05-23
Hugging FaceSLMLocal LLMParameters

The Best Small Language Models Available on Hugging Face

A comprehensive breakdown of the most capable sub-8B parameter local language models currently dominating the Hugging Face hub.

Read more →
guides
architect • 2026-05-23
Local LLMProjectsPythonAutomation

5 Cool Things I Did With Local Language Models

Discover five unique, practical, and incredibly fun projects you can build on consumer hardware using open source local LLMs.

Read more →
guides
architect • 2026-05-23
PythonSearchEmbeddingsLLMsNLP

Building Context-Aware Search in Python with LLM Embeddings and Metadata

A complete guide to constructing a semantic engine using SentenceTransformers and pre-filtering metadata to return context-rich results.

Read more →
guides
architect • 2026-05-23
Local LLMAgentsMachine LearningSLM

Top 5 Small Language Models for Agentic Tool Calling

Explore 5 small language models (SLMs) that feature open weights and first-class tool-calling support for local agentic workflows.

Read more →
News
architect • 2026-04-29T09:00:00Z
OpenAIAnthropicEnterprise AIAI TalentSoftware

The AI Talent War Comes for Enterprise Software

OpenAI and Anthropic are raiding the C-suites of Salesforce, Snowflake, and Palantir — not for researchers this time, but for seasoned enterprise sales and go-to-market executives.

Read more →
News
architect • 2026-04-28T10:00:00Z
MetaAgentsChinaM&ARegulation

China Blocks Meta's $2bn Acquisition of AI Start-up Manus

Beijing's top economic regulator has prohibited Meta's roughly $2 billion takeover of autonomous AI agent start-up Manus, citing foreign investment restrictions tied to the company's Chinese origins.

Read more →
Guide
architect • 2026-04-25T10:00:00Z
AgentsLocal LLMsSLMsOllamaLangChain

Building AI Agents with Local Small Language Models (SLMs)

Learn how to build fully functional, private AI agents on your own hardware using Ollama and LangChain with lightweight models under 10B parameters.

Read more →
News
architect • 2026-04-14T08:00:00Z
Multimodal AILLMsModelsReasoning

The Rise of Multi-Modal Reasoning in Next-Gen LLMs

How native multi-modality trained from the ground up is revolutionizing application development and changing the way we interact with AI.

Read more →
Article
architect • 2026-04-13T10:00:00Z
Vector DatabasesRAGEnterprise AIEmbeddings

Why Vector Databases are Essential for Enterprise AI

A deep dive into embeddings, semantic search, and why standard SQL databases aren't enough for LLM context retrieval.

Read more →