The welcome panel shown before a conversation has started.
Overview
A lightweight chat website that looks and feels like ChatGPT but runs entirely on your own machine using a locally hosted Ollama model - no data leaves the computer. A Flask server renders the chat page, saves every message to a plain-text file on disk so a conversation survives a reload, and streams the model's reply back word by word to feel more like a live response.
Architecture & Implementation
A Flask server renders the chat page from existing message history using server-side Jinja templates, falling back to a welcome panel when no conversation exists yet. Submitting a message hits an endpoint that validates and normalizes the input, appends it to a disk-backed conversation file, and only then invokes the local model.
The model call itself goes through LangChain and OllamaLLM, formatting the full exchange as a chat prompt with a fixed system prompt before sending it to the locally hosted model - nothing is sent anywhere outside the machine it runs on.
Conversation history is stored in a custom plain-text format rather than a database, structured so multi-line messages can be reconstructed losslessly on reload; parsing fails fast on an invalid file header rather than silently continuing with corrupted history. A separate endpoint resets the file back to its known initial state to clear the conversation.
On the frontend, messages submit on Enter (Shift+Enter for a newline), the model's reply is streamed back word by word to feel more like a live response, and the chat view auto-scrolls and auto-resizes its input as the conversation grows.