

Putting generative AI into mobile apps stopped being marketing gloss a long time ago. For a chatbot to deliver real value, cut customer support costs and never put brand reputation at risk through hallucinations, its architecture has to be designed with precision.
This case study walks through two radically different technical and architectural approaches we delivered for our clients: 2N, a global leader in access control systems, and FlexiFin, a fast-growing fintech.

RAG (Retrieval-Augmented Generation) semantic search engine over extensive technical documentation.
Speed up the work of installation technicians in the field who need instant technical advice without tediously browsing hundreds of PDF pages, and relieve 2N's overloaded internal technical support.





Hundreds of manuals converted into vector data. The chatbot understands a technician's question even if they don't use the exact technical term.
With technical devices, incorrect AI advice (e.g. wrong pin wiring) can destroy hardware worth tens of thousands of crowns. That's why the model doesn't answer from its general knowledge. It acts as an intelligent search engine - every answer is based on results from the embedding database of 2N manuals and backed by a source.
The chatbot is locked in software to 2N product technical support only. Attempts to steer the conversation to other topics are filtered out immediately.
The chatbot doesn't just answer - it takes the technician straight to the specific chapter of the manual integrated in the app (displayed via an optimized offline WebView). This is the feature users value most in practice. When they are troubleshooting and need information fast, they get directly to the relevant part of the source.
It runs on OpenAI gpt-4o, which offered the ideal balance of price, quality and ease of use. However, the app is ready to switch to a different model in the future if the terms of service or the client's needs change.
The Qdrant database enables fast storage and retrieval of embeddings.
We use the LangChain framework, which saved time compared to a custom solution.
We split manuals into chapters - smaller text blocks (chunks) - so the AI retains the context of tables and wiring diagrams.




















Model-agnostic orchestrator of standard APIs with strict EU data protection for a regulated fintech (Czech National Bank).
Reduce the call centre load (fewer calls about account status, repayments, etc.) and communicate new in-app features directly through natural language, without constantly building complex new UI.










The bot isn't tied to a single provider. The architecture allows LLM models to be swapped easily in the background (the pilot starts with standard OpenAI; production is planned to move to models with guaranteed EU hosting, e.g. Claude or local open-source models).
The bot is orchestrated to use tokens economically: it requests and retrieves information only at the moment it actually needs it. Instead of constantly sending the entire client context to the LLM (which would be extremely token-expensive), the bot first calls the app's internal standard APIs to fetch specific information (e.g. about cards or client data). For example, if a client asks the chatbot about their repayment amount, the bot requests only that piece of information and doesn't load other data about the user, the services they use and so on - keeping operation as token-sustainable as possible.
Strong emphasis on data sovereignty in a regulated fintech environment. The entire conversation history is stored and logged on servers in the EU for potential future audits.
Users can start a conversation directly from interactive banners in the app, which dramatically reduces friction (number of clicks).
The app is ready for future voice control of the chatbot (voice-control ready).
The database stores all data anonymously while enabling future analysis to improve the chatbot's answers and performance.
Users can have multiple chats running at once, all stored anonymously in the database. For clarity, and so the chatbot always shows the most up-to-date data and doesn't confuse users with older information from previous conversations (e.g. repayment amounts, which can change over time), users only see their most recent chats from the last few hours.
The bot's backend runs in the client's environment for maximum security and reliability.











Two completely different worlds. For 2N it was about absolute precision in technical search; for FlexiFin, about secure data orchestration.


We've shown two completely different worlds. For 2N it was about absolute precision in technical search; for FlexiFin, about secure data orchestration.
We don't clone one boxed solution for everyone. We always start by analysing your business, data and users. Only then do we build technology that makes sense – without unnecessary hype or wasted budget.
Got an idea for your own AI service?
Get in touchWe'd love to design a working, tailor-made solution for you.