# Automate Gmail with AI in 2026 (No Fluff, Just Code)
You get 200+ emails a day. Most are noise. Some require action. You don’t want to read them all—you want to *process* them.
AI can help. But most tutorials stop at “use Zapier + Gemini to auto-respond.” That’s not automation—it’s a band-aid.
Real automation means:
– Prioritizing urgent threads
– Extracting structured data (e.g., invoice amounts, ticket IDs)
– Drafting replies with context-aware tone
– Archiving or flagging based on *meaning*, not just keywords
In 2026, Gmail’s API + local LLMs (like Llama 3.2 or Phi-3) give you full control—no vendor lock-in, no per-email costs. Let’s build it.
—
## Understand Gmail’s Real Limits (Before You Code)
Gmail’s API has three layers you need to know:
| Layer | What It Does | Limitations |
|——-|————–|————-|
| **Gmail API (REST)** | Read/write emails, labels, threads | 250M requests/day free tier; rate-limited to 250/sec/user |
| **Google Apps Script (GAS)** | Triggers on new mail, run serverless functions | 90-min runtime/day; no long-running processes |
| **OAuth2 + Local Client** | Full control, runs on your machine/server | Requires manual auth flow; no built-in re-auth |
⚠️ **Critical Reality Check**:
Gmail API *does not* give you email *content* in one call. You must:
1. List thread IDs
2. Fetch messages by ID (with `format=metadata` or `full`)
3. Decode base64url-encoded body (for `text/plain` or `text/html`)
No library hides this cleanly. `google-api-python-client` is the official option, but it’s verbose. We’ll use it—because you need to know what’s really happening.
—
## Step 1: Set Up Your Environment (Zero Dependencies on Google Cloud)
You don’t need a billing-enabled Cloud project. Here’s the minimal setup:
1. Go to [Google Cloud Console > APIs & Services > Credentials](https://console.cloud.google.com/apis/credentials)
2. Click **+ Create Credentials** → **OAuth client ID**
3. Application type: **Desktop app**
4. Download `credentials.json` → save to `~/.gmail-ai/credentials.json`
5. Install Python deps:
“`bash
pip install google-api-python-client google-auth-httplib2 google-auth-oauthlib
“`
**No API key. No service account.** This is user-based auth—critical because Gmail only lets you act *as* the user (not *on behalf of* all users in a domain without Workspace admin).
—
## Step 2: Fetch Emails (With Real Code)
Here’s how to list unread threads and fetch their bodies:
“`python
from google.oauth2.credentials import Credentials
from googleapiclient.discovery import build
import base64
import json
def get_service():
creds = Credentials.from_authorized_user_file(‘credentials.json’, [‘https://www.googleapis.com/auth/gmail.modify’])
return build(‘gmail’, ‘v1′, credentials=creds)
def fetch_unread_emails(service, max_results=10):
# Step 1: Get thread IDs for unread messages
results = service.users().threads().list(
userId=’me’,
q=’is:unread’,
maxResults=max_results
).execute()
threads = results.get(‘threads’, [])
emails = []
for thread in threads:
# Step 2: Fetch full thread data
thread_data = service.users().threads().get(userId=’me’, id=thread[‘id’]).execute()
for msg in thread_data[‘messages’]:
payload = msg[‘payload’]
headers = {h[‘name’]: h[‘value’] for h in payload.get(‘headers’, [])}
# Skip duplicates in same thread
if ‘Subject’ not in headers:
continue
body = “”
if ‘parts’ in payload:
for part in payload[‘parts’]:
if part[‘mimeType’] == ‘text/plain’:
data = part[‘body’].get(‘data’, ”)
body = base64.urlsafe_b64decode(data).decode(‘utf-8’)
break
else: # Simple message
data = payload[‘body’].get(‘data’, ”)
body = base64.urlsafe_b64decode(data).decode(‘utf-8’)
emails.append({
‘thread_id’: thread[‘id’],
‘subject’: headers.get(‘Subject’, ”),
‘from’: headers.get(‘From’, ”),
‘body’: body.strip()
})
return emails
# Usage
service = get_service()
emails = fetch_unread_emails(service)
print(f”Fetched {len(emails)} emails”)
“`
**What’s happening**:
– `threads().list()` with `q=’is:unread’` avoids scanning all mail (saves quota)
– We decode `base64url` manually—no shortcuts. Gmail *always* base64-encodes body content.
– We stop at the first `text/plain` part. HTML parsing is overkill for most automation.
**Limitation**: This fetches *only* unread threads. If you need to process older emails, use `q=’after:2026/01/01’` (or any date filter).
—
## Step 3: Use a Local LLM to Process Emails (No API Keys)
You don’t need Gemini Pro or GPT-4 for this. For email triage, a 3B-parameter model is plenty. Let’s use `llama.cpp` with a quantized Llama 3.2-3B model.
### Why local?
– Zero per-email cost
– Full data control (emails never leave your machine)
– Works offline
**Install**:
“`bash
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp && make
“`
Download a quantized model (e.g., `Meta-Llama-3.2-3B-Instruct-Q4_K_M.gguf` from Hugging Face → ~2.1GB).
### Prompt Template for Email Triage
“`python
PROMPT = “””You are an email triage assistant. Classify the email and suggest next steps.
Email:
Subject: {subject}
From: {sender}
Body: {body}
Output ONLY valid JSON with these fields:
– priority: “high” | “medium” | “low”
– action: “reply” | “archive” | “flag” | “schedule”
– summary: “10-word max summary”
– details: “optional follow-up notes”
“””
“`
### Run Inference (Python)
“`python
import subprocess
import json
import os
def run_llama(prompt: str, model_path: str = “models/Meta-Llama-3.2-3B-Instruct-Q4_K_M.gguf”) -> dict:
# Llama.cpp expects system prompt + user prompt
full_prompt = f”<|begin_of_text|><|start_header_id|>system<|end_header_id|>\nYou are an email triage assistant.<|eot_id|><|start_header_id|>user<|end_header_id|>\n{prompt}<|eot_id|><|start_header_id|>assistant<|end_header_id|>\n”
# Run llama.cpp inference
result = subprocess.run(
[“./main”,
“-m”, model_path,
“-p”, full_prompt,
“-n”, “256”,
“–temp”, “0.1”,
“–repeat-penalty”, “1.1”],
capture_output=True,
text=True,
cwd=”llama.cpp”
)
# Extract JSON from output (after last “assistant” token)
raw_output = result.stdout.split(“<|assistant|>“)[-1].strip()
try:
return json.loads(raw_output)
except json.JSONDecodeError:
# Fallback: return raw output as summary
return {
“priority”: “medium”,
“action”: “flag”,
“summary”: raw_output[:100],
“details”: “”
}
# Process one email
email = emails[0]
prompt = PROMPT.format(
subject=email[‘subject’],
sender=email[‘from’],
body=email[‘body’][:1000] # Trim to avoid context overflow
)
result = run_llama(prompt)
print(json.dumps(result, indent=2))
“`
**Why this works**:
– Llama 3.2-3B handles structured output reliably at low temp (`0.1`)
– We trim body to 1000 chars—enough for triage, avoids context overflow
– Fallback handles model hallucinations (returns `flag` action if JSON fails)
**Performance**: ~12 seconds per email on a 16GB RAM Mac M2. Not real-time, but fine for batch processing.
—
## Step 4: Automate Actions (Label, Reply, Archive)
Gmail API lets you apply labels, but *not* send replies directly from `text/plain` content. You must:
1. Create a draft with `message/rfc822` MIME
2. Send the draft
### Example: Auto-flag high-priority emails
“`python
def apply_label(service, thread_id: str, label_name: str):
# Check if label exists
labels = service.users().labels().list(userId=’me’).execute()[‘labels’]
label_id = next((l[‘id’] for l in labels if l[‘name’] == label_name), None)
if not label_id:
# Create label if missing
label = service.users().labels().create(
userId=’me’,
body={‘name’: label_name, ‘labelListVisibility’: ‘labelShow’}
).execute()
label_id = label[‘id’]
# Apply label to thread
service.users().threads().modify(
userId=’me’,
id=thread_id,
body={‘addLabelIds’: [label_id]}
).execute()
# Apply “HIGH_PRIORITY” label to flagged emails
for email in emails:
if result[‘priority’] == ‘high’:
apply_label(service, email[‘thread_id’], ‘HIGH_PRIORITY’)
“`
### Draft a reply (with LLM-generated content)
“`python
def create_draft_reply(service, thread_id: str, reply_text: str, original_email: dict):
import email.mime.text
import base64
# Build MIME message
msg = email.mime.text.MIMEText(reply_text)
msg[‘to’] = original_email[‘from’]
msg[‘subject’] = f”Re: {original_email[‘subject’]}”
msg[‘in-reply-to’] = original_email[‘message_id’] # Must have this field!
raw = base64.urlsafe_b64encode(msg.as_bytes()).decode()
draft = service.users().drafts().create(
userId=’me’,
body={‘message’: {‘raw’: raw}}
).execute()
print(f”Draft created: {draft[‘id’]}”)
“`
⚠️ **Critical**: `in-reply-to` must match the original email’s `Message-ID` header. If you don’t set it, Gmail won’t thread the reply. We’ll need to fetch the `Message-ID` in our earlier `fetch_unread_emails` function.
—
## Step 5: Batch Processing (Don’t Get Rate-Limited)
You’ll hit the 250/sec limit fast.


