On-premises AI assistant on self-hosted LLMs
COMAN Software GmbH · AI Engineer · Fullstack Developer · 08/2025 – 04/2026
COMAN Software wanted to give its users secure, AI-powered access to product and company knowledge without violating data-protection requirements. The chatbot the company had been running on HubSpot was costly to operate.
I designed and built a Q&A chatbot powered entirely by locally hosted language models. The whole system runs on-premises: LM Studio and Ollama serve the models on an NVIDIA DGX Spark, Qdrant handles vector search, and PostgreSQL stores the application data. No sensitive data ever leaves the company's own infrastructure for an external cloud service. I tracked answer quality and latency through Langfuse traces and improved them iteratively from there.
What I built
- RAG pipeline indexed over COMAN's product and company knowledge
- Integration with the HubSpot API
- Direct integration into COMAN's own software products
- Containerized deployment via Docker
Result: the costly HubSpot chatbot was fully replaced. Costs are now predictable and significantly lower, while answer quality is noticeably better, with far fewer hallucinations. The chatbot is today an integral part of the COMAN software.
Planning an AI project?
Free first assessment: I'll tell you whether agents are worth it for your use case — or whether a simpler system will do.