Writing and Content
Text generation, copywriting, translation, and content pipelines with local models.
New to AI? basics/ β what models are, how to choose and set them up.
π·πΊ Russian version: writing.ru.md
β Use cases Β· Reflection pattern β
- Recommended models
- Quantization for text
- General text generation
- Reflection pattern
- Mass production
- RAG for content
- Whats next
1. Recommended models
| Task | Model | Run | tok/s | Why |
|---|---|---|---|---|
| Quick drafts | Qwen 3.5 4B | ollama run qwen3.5:4b |
28β35 | Good quality, fast |
| Quality texts | Qwen 3.5 9B | ollama run qwen3.5:9b |
10β13 | Best balance on 16 GB |
| Classic | Llama 3.1 8B | ollama run llama3.1:8b |
14β18 | Stable style |
| Analytics / reports | Phi-4-mini | ollama run phi4-mini |
25β30 | Best for reasoning |
2. Quantization for text
For writing and analytics, Q5_K_M makes sense over Q4_K_M β if the model fits in RAM:
| Format | When to use | Difference |
|---|---|---|
| Q4_K_M | Default, drafts | Fast, 3.5% quality loss |
| Q5_K_M | Final texts, analytics | 2.5% loss, slightly more memory |
| Q6_K | Legal / medical wording | 1.6% loss, noticeably better |
# Example: run model with Q5_K_M
ollama pull qwen3.5:9b:q5_k_m
ollama run qwen3.5:9b:q5_k_m
More details β local-models/quantization.md.
3. General text generation
import requests
def generate(model, prompt):
r = requests.post("http://localhost:11434/api/generate", json={
"model": model, "prompt": prompt, "stream": False
})
return r.json()["response"]
for task, model in [
("Write a tweet about AI", "qwen3.5:4b"),
("Write a technical article about RAG", "qwen3.5:9b"),
]:
print(f"Using {model}...")
4. Reflection pattern
Generate β critique β improve. A single pass gives a βrawβ result; reflection adds self-review:
Draft (Qwen 4B, fast) β Reviewer agent (Qwen 9B + Reflection) β Final
Here is a compact single-function implementation:
def reflection_article(topic):
r1 = requests.post("http://localhost:11434/api/chat", json={
"model": "qwen3.5:9b",
"messages": [{"role": "user", "content": f"Write an article about: {topic}"}]
})
draft = r1.json()["message"]["content"]
r2 = requests.post("http://localhost:11434/api/chat", json={
"model": "qwen3.5:4b",
"messages": [
{"role": "system", "content": "You are a strict editor. Find issues."},
{"role": "user", "content": draft}
]
})
feedback = r2.json()["message"]["content"]
r3 = requests.post("http://localhost:11434/api/chat", json={
"model": "qwen3.5:9b",
"messages": [
{"role": "system", "content": "Improve based on the feedback."},
{"role": "user", "content": f"Draft:\n{draft}\n\nFeedback:\n{feedback}\n\nImproved:"}
]
})
return r3.json()["message"]["content"]
Alternatively, a modular approach with separate functions:
import requests
OLLAMA = "http://localhost:11434/api/chat"
def draft(text: str) -> str:
"""Quick draft."""
r = requests.post(OLLAMA, json={
"model": "qwen3.5:4b",
"messages": [{"role": "user", "content": text}],
"stream": False
})
return r.json()["message"]["content"]
def review(text: str) -> str:
"""Critical review."""
r = requests.post(OLLAMA, json={
"model": "qwen3.5:9b",
"messages": [
{"role": "system", "content": (
"You are an editor. Find errors, inaccuracies, weak spots in the text. "
"Check: facts, grammar, style, structure."
)},
{"role": "user", "content": text}
],
"stream": False
})
return r.json()["message"]["content"]
def improve(text: str, critique: str) -> str:
"""Improve based on critique."""
r = requests.post(OLLAMA, json={
"model": "qwen3.5:9b",
"messages": [
{"role": "system", "content": "Improve the text based on the critique. Return only the final version."},
{"role": "user", "content": f"Original:\n{text}\n\nCritique:\n{critique}\n\nImproved text:"}
],
"stream": False
})
return r.json()["message"]["content"]
# Full pipeline
def write_with_reflection(topic: str) -> str:
"""Writes text with self-review."""
print(" βοΈ Draft...")
raw = draft(f"Write a short article on: {topic}")
print(" π Review...")
critique = review(raw)
print(" β¨ Final version...")
final = improve(raw, critique)
return final
# Example
article = write_with_reflection("Advantages of local AI models")
print(article)
Result: the text goes through three stages β generation β critique β improvement. Quality is noticeably higher than a single pass.
5. Mass production
Generate from CSV with configurable tone:
#!/usr/bin/env python3
"""content_generator.py β mass text generation through Ollama."""
import requests
import csv
from pathlib import Path
OLLAMA = "http://localhost:11434/api/chat"
def generate(topic: str, tone: str = "neutral") -> str:
"""Generates text by topic with given tone."""
tones = {
"neutral": "Write an informative text.",
"professional": "Write a business text. Use professional vocabulary.",
"friendly": "Write a friendly, conversational text.",
"persuasive": "Write a persuasive text with a call to action."
}
response = requests.post(OLLAMA, json={
"model": "qwen3.5:4b",
"messages": [
{"role": "system", "content": tones.get(tone, tones["neutral"])},
{"role": "user", "content": f"Topic: {topic}"}
],
"stream": False
})
return response.json()["message"]["content"]
# Generate from CSV file
def generate_from_csv(csv_path: str, output_dir: str = "output"):
"""Generates texts for each CSV row."""
Path(output_dir).mkdir(exist_ok=True)
with open(csv_path) as f:
reader = csv.DictReader(f)
for row in reader:
text = generate(row["topic"], row.get("tone", "neutral"))
filename = row.get("filename", f"{row['topic'][:30]}.md")
filepath = Path(output_dir) / filename
filepath.write_text(text)
print(f" β {filename}")
if __name__ == "__main__":
# Example: mass generate descriptions
topics = [
{"topic": "What is RAG in the context of LLM", "tone": "professional"},
{"topic": "How to choose a model for coding", "tone": "friendly"},
{"topic": "Advantages of local AI over cloud", "tone": "persuasive"},
]
for item in topics:
text = generate(item["topic"], item["tone"])
print(f"\nπ {item['topic']}")
print(text[:200] + "...")
Also keep the simpler loop approach:
topics = ["What is RAG", "Local LLM setup", "AI agents explained"]
for topic in topics:
article = reflection_article(topic)
filename = topic.lower().replace(" ", "_") + ".md"
with open(filename, "w") as f:
f.write(f"# {topic}\n\n{article}")
6. RAG for content
For texts based on your own materials (documentation, knowledge base) β use RAG:
Topic β Search similar materials in database β Context + query β Text
More details β rag.md.
7. Whats next
| If you want | Go to |
|---|---|
| Use the Reflection agent for content | patterns.md |
| Build a content team of agents | tutorials/02-agent-team.md |
| Choose models for text | local-models/catalog.md |
| Automate content pipelines | automation.md |
| Back to use cases | README.md |
In section: coding Β· rag Β· automation Β· writing
Related sections: Local Models Β· AI Agents Β· Zero Level
Navigation: β Use Cases Β· β Back to main Β· π·πΊ Π ΡΡΡΠΊΠΈΠΉ