What you can actually build with a Raspberry Pi and AI
A Raspberry Pi can run a small AI assistant that listens to voice commands, answers questions, and controls smart home devices — but it will not work like Alexa or Siri. The assistant runs only on your local network, responds more slowly, and needs you to set up each piece yourself. What you get is a learning project that teaches you how AI assistants work, plus a device you control completely without sending data to a company's servers.
The most realistic goal is a voice assistant that recognizes a wake word (like "Hey Pi"), records what you say, sends it to a speech-to-text service, processes the text through an AI model, and reads the answer aloud. You will need a Raspberry Pi 4 or 5 (the Pi 3 is too slow), a microphone, a speaker, and about two to four hours to get it working.
Key Takeaways
- A Raspberry Pi 4 or 5 with at least 4GB of RAM is the minimum; the Pi 5 with 8GB handles more tasks without slowing down.
- You will need a USB microphone and speaker, a microSD card with at least 32GB, and a power supply rated for at least 3 amps.
- Most working setups use Ollama (a local AI model runner), Whisper (speech-to-text), and a text-to-speech tool like pyttsx3 or Festival.
- The whole project costs between $80 and $150 in hardware, plus the time to install and configure software on the command line.
- Your assistant will be slower and less capable than cloud-based ones, but it runs entirely offline and you own the data.
Hardware you need to start
Start with a Raspberry Pi 4 Model B with 4GB RAM minimum — the 8GB version is better if you plan to run larger AI models. The Pi 5 is faster but costs more. Do not use a Pi 3 or earlier; they lack the processing power and RAM to run modern AI models without freezing.
You also need a microSD card rated for at least 32GB (a 64GB card gives you room to experiment). Buy one marked "high endurance" or "A2" rated, because a slow card will make the whole system feel sluggish. A power supply rated for at least 3 amps at 5 volts is essential — the official Raspberry Pi power supply works, but many cheap ones do not deliver enough power under load.
For audio input and output, buy a USB microphone (around $15 to $30) and a USB speaker or 3.5mm speaker (around $20 to $50). Avoid tiny USB microphones designed for video calls; they pick up too much background noise. A basic USB condenser microphone or a webcam with a built-in mic works better. Test both before you start the software setup.
Total hardware cost is typically $100 to $150 if you already have a monitor, keyboard, and mouse to set up the Pi initially. You can use a laptop or desktop to do the setup over SSH (remote command line) if you do not have spare peripherals.
Installing the operating system and dependencies
Download Raspberry Pi OS Lite (the command-line version, not the desktop version) from the official Raspberry Pi website. Use the Raspberry Pi Imager tool to write it to your microSD card. Lite uses less RAM and CPU, leaving more resources for AI tasks.
Insert the card into the Pi, connect power, and wait about two minutes for it to boot. Connect to it over SSH from another computer on your network, or plug in a monitor and keyboard. Update the system by running these commands one at a time:
- sudo apt update
- sudo apt upgrade
- sudo apt install python3-pip python3-venv git
This installs Python (the language you will use to write the assistant), pip (the package manager), and Git (to download code from repositories). The upgrade step takes 10 to 15 minutes on a Pi 4.
Next, install audio tools so the Pi can record and play sound:
- sudo apt install alsa-utils pulseaudio
- sudo usermod -a -G audio pi
Reboot the Pi after this step. Test your microphone and speaker by running alsamixer to adjust volume levels, then record a test sound with arecord -d 5 test.wav and play it back with aplay test.wav.
Setting up Ollama for local AI models
Ollama is software that downloads and runs AI language models on your Raspberry Pi without needing an internet connection after setup. It handles the heavy lifting of running the model efficiently on limited hardware.
Install Ollama by running:
- curl https://ollama.ai/install.sh | sh
- ollama serve (this starts the Ollama service)
In a new terminal window (or SSH session), download a small model that fits on a Pi:
- ollama pull mistral (downloads the Mistral model, about 4GB)
The download takes 10 to 20 minutes depending on your internet speed. Once it finishes, test it by running ollama run mistral and typing a question. Type /bye to exit.
If 4GB is too large for your storage, use ollama pull neural-chat instead (about 2GB). The smaller model is less capable but runs faster on a Pi 4. Keep Ollama running in the background; you will connect to it from your Python code.
Adding speech-to-text and text-to-speech
Whisper is OpenAI's speech-to-text tool that runs locally on your Pi. Install it with:
- pip3 install openai-whisper
Download the small model (about 140MB) by running whisper --model tiny --language en --task transcribe test.wav. This downloads the model and transcribes a test audio file. The "tiny" model is fast enough for a Pi; larger models are more accurate but much slower.
For text-to-speech (reading answers aloud), install pyttsx3, which works offline:
- pip3 install pyttsx3
Test it by creating a file called test_tts.py with this code:
import pyttsx3 engine = pyttsx3.init() engine.say("Hello, I am your Raspberry Pi assistant") engine.runAndWait()
Run it with python3 test_tts.py. You should hear the text spoken aloud through your speaker. If you do not hear anything, check that your speaker is plugged in and the volume is turned up in alsamixer.
Writing the assistant code
Create a new Python file called assistant.py. This is the main program that ties everything together. Here is a basic version that listens for a wake word, records audio, transcribes it, sends it to the AI model, and reads the response:
import subprocess import whisper import pyttsx3 import requests import json engine = pyttsx3.init() model = whisper.load_model("tiny") def listen_for_audio(duration=5): subprocess.run(["arecord", "-d", str(duration), "input.wav"]) return "input.wav" def transcribe_audio(audio_file): result = model.transcribe(audio_file) return result["text"] def query_ollama(prompt): response = requests.post("http://localhost:11434/api/generate", json={"model": "mistral", "prompt": prompt, "stream": False}) return response.json()["response"] def speak(text): engine.say(text) engine.runAndWait() print("Listening...") audio = listen_for_audio(5) text = transcribe_audio(audio) print(f"You said: {text}") response = query_ollama(text) print(f"Assistant: {response}") speak(response)
Save this file and run it with python3 assistant.py. Speak clearly into your microphone for 5 seconds. The program will transcribe what you said, send it to Ollama, and read the response aloud. The first run takes longer because Whisper and Ollama load their models into memory.
This basic version works but does not listen for a wake word. To add that, you would need to install a wake-word detector like Porcupine or PocketSphinx, which adds complexity. Start with this version and add features once you understand how each piece works.
Troubleshooting common problems
The microphone is not recording. Run arecord -l to list audio devices. If your USB microphone does not appear, unplug it, wait 10 seconds, and plug it back in. Then run sudo alsamixer and make sure the microphone input is not muted (look for "MM" which means muted; press M to unmute). Check that the microphone is selected as the default input device.
Whisper is very slow or crashes. The "tiny" model is the smallest; if it still crashes, your Pi may not have enough free RAM. Close other programs and try again. If you have a Pi 4 with only 2GB RAM, you may need to add a swap file by running sudo dphys-swapfile swapon, though this slows everything down significantly.
Ollama is not responding. Make sure the Ollama service is still running. In a separate terminal, run ollama serve again. If it says the port is already in use, Ollama is running but may be stuck. Restart it with sudo systemctl restart ollama (if you set it up as a service) or kill the process and start it again.
The assistant takes 30 seconds to answer. This is normal on a Pi 4. The Whisper model takes 5 to 10 seconds to transcribe, and Ollama takes 10 to 20 seconds to generate a response, depending on how long the answer is. A Pi 5 is noticeably faster. If speed matters, consider using a cloud API for transcription or response generation, though that requires an internet connection and sends your data to a company's servers.
Next steps and improvements
Once the basic assistant works, you can add features. A wake-word detector lets the Pi listen continuously and only record when it hears "Hey Pi" or another phrase you choose. Libraries like Porcupine (paid) or PocketSphinx (free) do this, though they add CPU load.
You can also connect the assistant to smart home devices. If you have Philips Hue lights or a smart plug, add code that parses commands like "turn on the bedroom light" and sends the right signal to your devices. This requires learning the API for each device, but it makes the assistant actually useful.
Another direction is to run a larger AI model. The Mistral model is small, but Ollama also supports Llama 2, Neural Chat, and others. Larger models give better answers but need more RAM and storage. Experiment with different models to find the balance between quality and speed on your hardware.
Finally, consider running the assistant as a service that starts automatically when the Pi boots. This turns it into a permanent device rather than a project you run manually. Use systemd to create a service file that launches your Python script on startup.
Frequently Asked Questions
Can I use a Raspberry Pi 3 for this?
A Pi 3 has only 1GB of RAM and a much slower processor. Whisper and Ollama will run, but responses take several minutes and the system may freeze. A Pi 4 with 4GB is the practical minimum. If you already have a Pi 3, use it to learn the software setup, then upgrade to a Pi 4 or 5 for a usable assistant.
Do I need an internet connection?
You need internet to download Ollama, Whisper, and the AI models. After that, the assistant runs entirely offline. No audio or text leaves your network. This is different from Alexa or Google Assistant, which send everything to company servers.
Can I make this assistant control my smart home?
Yes, but it requires additional code. You would parse the text response from Ollama, look for keywords like "light" or "temperature", and send commands to your smart devices. This works best with devices that have open APIs, like Philips Hue or Home Assistant. Proprietary systems like Amazon Alexa are harder to integrate.
How much does this cost compared to buying Alexa?
A Raspberry Pi setup costs $100 to $150 upfront. An Echo Dot costs $30 to $50. However, the Pi is a one-time purchase that you own completely, while Alexa requires a subscription for some features and sends data to Amazon. The Pi is more expensive but gives you full control and privacy.
What if I want to use a cloud AI service instead of Ollama?
You can replace the Ollama code with calls to OpenAI's API, Google's Vertex AI, or another service. This makes responses faster and more accurate, but costs money per request and requires an internet connection. For learning, Ollama is free and teaches you how the pieces fit together.