Skip to content

Ollama

Overview

  • Run LLMs Locally: It lets you download and run large language models directly on your own Mac, Windows, or Linux machine. Uses your computer's own processing power (CPU or GPU)
  • Simplified Setup: Provides a easy-to-install package.
  • Privacy and Offline Use: Since models run on your hardware, your conversations are private and don't leave your device. It also works without an internet connection after the initial model download.
  • Command-Line Interface: Provides simple command-line chat interface through the terminal (e.g., ollama run llama3.2).
  • Developer-Friendly API: Ollama automatically creates a local REST API, making it easy for developers to build applications powered by local LLMs.
  • Model Library: Provides access to a growing library of open-source models.
  • Customization: Allows you to create and customize your own model variants using a simple configuration file called a Modelfile.
  • Model Offloading: Automatically unloads models from memory when not in use.
  • DOCS: https://github.com/ollama/ollama/tree/main/docs

Up and Running

# NOTE: Add `--gpus=all` if you have GPUs
podman run -d \
  --name ollama \
  --network llm-stack \
  -v ollama:/root/.ollama \
  -p 127.0.0.1:11434:11434 \
  ollama/ollama:0.9.6

podman logs -f ollama
podman exec -ti ollama bash
  ollama pull llama3.2:1b
  ollama list
  ollama run llama3.2:1b
    What are some fun things to do in columbus, ohio
  ollama pull bge-m3:567m
  exit

Usage

curl http://localhost:11434/v1/models | jq
curl http://localhost:11434/v1/chat/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "llama3.2:1b",
        "messages": [
            {
                "role": "system",
                "content": "You are a helpful assistant."
            },
            {
                "role": "user",
                "content": "Hello!"
            }
        ]
    }' | jq