Ollama
Overview
- Run LLMs Locally: It lets you download and run large language models directly on your own Mac, Windows, or Linux machine. Uses your computer's own processing power (CPU or GPU)
- Simplified Setup: Provides a easy-to-install package.
- Privacy and Offline Use: Since models run on your hardware, your conversations are private and don't leave your device. It also works without an internet connection after the initial model download.
- Command-Line Interface: Provides simple command-line chat interface through the terminal (e.g.,
ollama run llama3.2).
- Developer-Friendly API: Ollama automatically creates a local REST API, making it easy for developers to build applications powered by local LLMs.
- Model Library: Provides access to a growing library of open-source models.
- Customization: Allows you to create and customize your own model variants using a simple configuration file called a
Modelfile.
- Model Offloading: Automatically unloads models from memory when not in use.
- DOCS: https://github.com/ollama/ollama/tree/main/docs
Up and Running
# NOTE: Add `--gpus=all` if you have GPUs
podman run -d \
--name ollama \
--network llm-stack \
-v ollama:/root/.ollama \
-p 127.0.0.1:11434:11434 \
ollama/ollama:0.9.6
podman logs -f ollama
podman exec -ti ollama bash
ollama pull llama3.2:1b
ollama list
ollama run llama3.2:1b
What are some fun things to do in columbus, ohio
ollama pull bge-m3:567m
exit
Usage
curl http://localhost:11434/v1/models | jq
curl http://localhost:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "llama3.2:1b",
"messages": [
{
"role": "system",
"content": "You are a helpful assistant."
},
{
"role": "user",
"content": "Hello!"
}
]
}' | jq