Writing and Content

Text generation, copywriting, translation, and content pipelines with local models.

New to AI? basics/ β€” what models are, how to choose and set them up.

πŸ‡·πŸ‡Ί Russian version: writing.ru.md


← Use cases Β· Reflection pattern β†’


  1. Recommended models
  2. Quantization for text
  3. General text generation
  4. Reflection pattern
  5. Mass production
  6. RAG for content
  7. Whats next

Task Model Run tok/s Why
Quick drafts Qwen 3.5 4B ollama run qwen3.5:4b 28–35 Good quality, fast
Quality texts Qwen 3.5 9B ollama run qwen3.5:9b 10–13 Best balance on 16 GB
Classic Llama 3.1 8B ollama run llama3.1:8b 14–18 Stable style
Analytics / reports Phi-4-mini ollama run phi4-mini 25–30 Best for reasoning

2. Quantization for text

For writing and analytics, Q5_K_M makes sense over Q4_K_M β€” if the model fits in RAM:

Format When to use Difference
Q4_K_M Default, drafts Fast, 3.5% quality loss
Q5_K_M Final texts, analytics 2.5% loss, slightly more memory
Q6_K Legal / medical wording 1.6% loss, noticeably better
# Example: run model with Q5_K_M
ollama pull qwen3.5:9b:q5_k_m
ollama run qwen3.5:9b:q5_k_m

More details β€” local-models/quantization.md.


3. General text generation

import requests

def generate(model, prompt):
    r = requests.post("http://localhost:11434/api/generate", json={
        "model": model, "prompt": prompt, "stream": False
    })
    return r.json()["response"]

for task, model in [
    ("Write a tweet about AI", "qwen3.5:4b"),
    ("Write a technical article about RAG", "qwen3.5:9b"),
]:
    print(f"Using {model}...")

4. Reflection pattern

Generate β†’ critique β†’ improve. A single pass gives a β€œraw” result; reflection adds self-review:

Draft (Qwen 4B, fast) β†’ Reviewer agent (Qwen 9B + Reflection) β†’ Final

Here is a compact single-function implementation:

def reflection_article(topic):
    r1 = requests.post("http://localhost:11434/api/chat", json={
        "model": "qwen3.5:9b",
        "messages": [{"role": "user", "content": f"Write an article about: {topic}"}]
    })
    draft = r1.json()["message"]["content"]

    r2 = requests.post("http://localhost:11434/api/chat", json={
        "model": "qwen3.5:4b",
        "messages": [
            {"role": "system", "content": "You are a strict editor. Find issues."},
            {"role": "user", "content": draft}
        ]
    })
    feedback = r2.json()["message"]["content"]

    r3 = requests.post("http://localhost:11434/api/chat", json={
        "model": "qwen3.5:9b",
        "messages": [
            {"role": "system", "content": "Improve based on the feedback."},
            {"role": "user", "content": f"Draft:\n{draft}\n\nFeedback:\n{feedback}\n\nImproved:"}
        ]
    })
    return r3.json()["message"]["content"]

Alternatively, a modular approach with separate functions:

import requests

OLLAMA = "http://localhost:11434/api/chat"

def draft(text: str) -> str:
    """Quick draft."""
    r = requests.post(OLLAMA, json={
        "model": "qwen3.5:4b",
        "messages": [{"role": "user", "content": text}],
        "stream": False
    })
    return r.json()["message"]["content"]

def review(text: str) -> str:
    """Critical review."""
    r = requests.post(OLLAMA, json={
        "model": "qwen3.5:9b",
        "messages": [
            {"role": "system", "content": (
                "You are an editor. Find errors, inaccuracies, weak spots in the text. "
                "Check: facts, grammar, style, structure."
            )},
            {"role": "user", "content": text}
        ],
        "stream": False
    })
    return r.json()["message"]["content"]

def improve(text: str, critique: str) -> str:
    """Improve based on critique."""
    r = requests.post(OLLAMA, json={
        "model": "qwen3.5:9b",
        "messages": [
            {"role": "system", "content": "Improve the text based on the critique. Return only the final version."},
            {"role": "user", "content": f"Original:\n{text}\n\nCritique:\n{critique}\n\nImproved text:"}
        ],
        "stream": False
    })
    return r.json()["message"]["content"]

# Full pipeline
def write_with_reflection(topic: str) -> str:
    """Writes text with self-review."""
    print("  ✏️ Draft...")
    raw = draft(f"Write a short article on: {topic}")
    print("  πŸ” Review...")
    critique = review(raw)
    print("  ✨ Final version...")
    final = improve(raw, critique)
    return final

# Example
article = write_with_reflection("Advantages of local AI models")
print(article)

Result: the text goes through three stages β€” generation β†’ critique β†’ improvement. Quality is noticeably higher than a single pass.


5. Mass production

Generate from CSV with configurable tone:

#!/usr/bin/env python3
"""content_generator.py β€” mass text generation through Ollama."""

import requests
import csv
from pathlib import Path

OLLAMA = "http://localhost:11434/api/chat"

def generate(topic: str, tone: str = "neutral") -> str:
    """Generates text by topic with given tone."""
    
    tones = {
        "neutral": "Write an informative text.",
        "professional": "Write a business text. Use professional vocabulary.",
        "friendly": "Write a friendly, conversational text.",
        "persuasive": "Write a persuasive text with a call to action."
    }
    
    response = requests.post(OLLAMA, json={
        "model": "qwen3.5:4b",
        "messages": [
            {"role": "system", "content": tones.get(tone, tones["neutral"])},
            {"role": "user", "content": f"Topic: {topic}"}
        ],
        "stream": False
    })
    return response.json()["message"]["content"]


# Generate from CSV file
def generate_from_csv(csv_path: str, output_dir: str = "output"):
    """Generates texts for each CSV row."""
    
    Path(output_dir).mkdir(exist_ok=True)
    
    with open(csv_path) as f:
        reader = csv.DictReader(f)
        for row in reader:
            text = generate(row["topic"], row.get("tone", "neutral"))
            
            filename = row.get("filename", f"{row['topic'][:30]}.md")
            filepath = Path(output_dir) / filename
            filepath.write_text(text)
            print(f"  βœ“ {filename}")


if __name__ == "__main__":
    # Example: mass generate descriptions
    topics = [
        {"topic": "What is RAG in the context of LLM", "tone": "professional"},
        {"topic": "How to choose a model for coding", "tone": "friendly"},
        {"topic": "Advantages of local AI over cloud", "tone": "persuasive"},
    ]
    
    for item in topics:
        text = generate(item["topic"], item["tone"])
        print(f"\nπŸ“„ {item['topic']}")
        print(text[:200] + "...")

Also keep the simpler loop approach:

topics = ["What is RAG", "Local LLM setup", "AI agents explained"]
for topic in topics:
    article = reflection_article(topic)
    filename = topic.lower().replace(" ", "_") + ".md"
    with open(filename, "w") as f:
        f.write(f"# {topic}\n\n{article}")

6. RAG for content

For texts based on your own materials (documentation, knowledge base) β€” use RAG:

Topic β†’ Search similar materials in database β†’ Context + query β†’ Text

More details β€” rag.md.


7. Whats next

If you want Go to
Use the Reflection agent for content patterns.md
Build a content team of agents tutorials/02-agent-team.md
Choose models for text local-models/catalog.md
Automate content pipelines automation.md
Back to use cases README.md

In section: coding Β· rag Β· automation Β· writing
Related sections: Local Models Β· AI Agents Β· Zero Level
Navigation: ← Use Cases Β· ↑ Back to main Β· πŸ‡·πŸ‡Ί Русский