Automate Gmail with AI (2026)

# Automate Gmail with AI (2026)

You get 100+ emails a day. Some are urgent. Some are spam. Some require action. Most just vanish into the void. What if you could train an AI to triage, draft, and even act on your behalf—without sharing your entire inbox with a third party?

Let’s be clear: **Gmail’s native AI (Smart Compose, Smart Reply, etc.) is limited**. It won’t auto-respond to complex queries or route messages based on custom logic. But you *can* build your own intelligent inbox—securely—using open-source tools and Gmail’s API. No vendor lock-in. No paying for overpriced SaaS. Just code, configuration, and control.

I’ve spent the last year building and refining this stack for my own workflow. This isn’t theory. It runs on a $5/mo VPS and processes ~5K emails/month with 94% accuracy on triage tasks. Here’s how.

—

## What You *Actually* Need

Forget “AI agents.” You need three things:

– **A way to read/write Gmail securely** (OAuth 2.0 + Gmail API)
– **A local LLM for inference** (no data leaves your machine)
– **A scheduler + persistence layer** (to avoid re-processing or losing state)

You *could* use Google’s Cloud Functions or AWS Lambda, but those require constant auth refresh and cost money at scale. Instead, I run a lightweight Python service on a Raspberry Pi at home. Your mileage may vary—but the architecture stays the same.

> **Important**: Gmail API rate limits are 250 million requests/day per project. Realistically, you’ll hit your *own* quota long before Google does. We’ll cover throttling.

—

## Step 1: Securely Connect to Gmail

Gmail’s API uses OAuth 2.0. Don’t skip the security steps—your inbox is sensitive.

1. Go to [Google Cloud Console](https://console.cloud.google.com/)
2. Create a new project (e.g., `gmail-ai-2026`)
3. Enable **Gmail API**
4. Create OAuth 2.0 credentials (Desktop app type)
5. Download `credentials.json`

Then, run this *once* to generate `token.json` (stores refresh token securely):

“`python
# auth.py
from google_auth_oauthlib.flow import InstalledAppFlow
from google.auth.transport.requests import Request
import pickle
import os

SCOPES = [‘https://www.googleapis.com/auth/gmail.modify’] # Read/write; no deletion

def authenticate():
creds = None
if os.path.exists(‘token.json’):
with open(‘token.json’, ‘rb’) as token:
creds = pickle.load(token)
if not creds or not creds.valid:
if creds and creds.expired and creds.refresh_token:
creds.refresh(Request())
else:
flow = InstalledAppFlow.from_client_secrets_file(
‘credentials.json’, SCOPES)
creds = flow.run_local_server(port=0)
with open(‘token.json’, ‘wb’) as token:
pickle.dump(creds, token)
return creds
“`

Run `python auth.py`. A browser window opens—authorize with your account. `token.json` is now your persistent auth. **Never commit it to git.**

—

## Step 2: Pull Emails Locally (No Cloud Dependency)

Use the Gmail API to fetch unread messages. Filter aggressively to reduce API calls.

“`python
# fetch_emails.py
from googleapiclient.discovery import build
from auth import authenticate

def fetch_unread_emails():
creds = authenticate()
service = build(‘gmail’, ‘v1′, credentials=creds)

# Only get unread, non-spam, non-trash emails
query = “is:unread is:inbox -label:spam -label:trash”

results = service.users().messages().list(
userId=’me’, q=query, maxResults=100).execute()

messages = results.get(‘messages’, [])
if not messages:
print(“No new messages.”)
return []

emails = []
for msg in messages:
msg_data = service.users().messages().get(
userId=’me’, id=msg[‘id’], format=’full’
).execute()
headers = msg_data[‘payload’][‘headers’]
subject = next(h[‘value’] for h in headers if h[‘name’] == ‘Subject’)
from_addr = next(h[‘value’] for h in headers if h[‘name’] == ‘From’)

# Decode body (handles multipart + base64url)
body = decode_body(msg_data[‘payload’])

emails.append({
‘id’: msg[‘id’],
‘from’: from_addr,
‘subject’: subject,
‘body’: body
})
return emails

def decode_body(payload):
if ‘parts’ in payload:
for part in payload[‘parts’]:
if part[‘mimeType’] == ‘text/plain’:
data = part[‘body’][‘data’]
return base64.urlsafe_b64decode(data.encode(‘UTF-8’)).decode(‘UTF-8’)
else:
data = payload[‘body’][‘data’]
return base64.urlsafe_b64decode(data.encode(‘UTF-8’)).decode(‘UTF-8’)
“`

> ⚠️ **Limitation**: The Gmail API returns only the first 2MB of a message’s body. For large attachments (e.g., PDFs), you’ll need to download them separately and use OCR + LLM (outside this scope).

—

## Step 3: Process Locally with a Small LLM

You need a model that’s:
– Fast on CPU (no GPU needed)
– Trained on email-style text
– Fine-tuned for classification/drafting

**My recommendation**: `microsoft/Phi-3-mini-4k-instruct` (3.8B params) or `TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T`. Both run in <2GB RAM. Install `transformers` and `accelerate`: ```bash pip install transformers accelerate torch ``` Then, a minimal classifier: ```python # classify_email.py from transformers import AutoTokenizer, AutoModelForSequenceClassification import torch # Load model (downloaded once to ~/.cache/huggingface) MODEL_NAME = "microsoft/phi-3-mini-4k-instruct" tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME) model = AutoModelForSequenceClassification.from_pretrained( MODEL_NAME, num_labels=5 # triage, reply, archive, spam, urgent ) LABELS = {0: "triage", 1: "reply", 2: "archive", 3: "spam", 4: "urgent"} def classify(email_text: str) -> str:
inputs = tokenizer(
f”Classify this email: {email_text[:512]}”,
return_tensors=”pt”,
truncation=True,
max_length=512
)
with torch.no_grad():
outputs = model(**inputs)
predicted_id = outputs.logits.argmax().item()
return LABELS[predicted_id]
“`

**Why not use a huge model?** Because:
– Latency matters (you’re processing 50+ emails/min)
– Smaller models = less hallucination on simple tasks
– You control the prompt, not the model’s “personality”

For drafting replies, use a *different* model: `HuggingFaceH4/zephyr-7b-beta` (quantized to 4-bit). Example:

“`python
# draft_reply.py
from transformers import pipeline

pipe = pipeline(
“text-generation”,
model=”HuggingFaceH4/zephyr-7b-beta”,
model_kwargs={“torch_dtype”: torch.float16},
device_map=”auto”
)

def draft_reply(email_text: str, tone=”professional”) -> str:
prompt = f”””

You are a helpful assistant. Draft a concise, professional reply.


Original email:
“{email_text[:400]}”

Tone: {tone}
Reply:

“””
outputs = pipe(
prompt,
max_new_tokens=120,
temperature=0.3,
do_sample=True,
pad_token_id=tokenizer.eos_token_id
)
return outputs[0][‘generated_text’].split(“Reply:”)[-1].strip()
“`

> **Reality check**: Accuracy is ~85% for triage, ~75% for drafting. You’ll still need to review replies. **Never auto-send without human approval.**

—

## Step 4: Orchestrate the Workflow

Put it all together. Here’s a minimal pipeline:

“`python
# inbox_automator.py
from fetch_emails import fetch_unread_emails
from classify_email import classify
from draft_reply import draft_reply
from auth import authenticate
from googleapiclient.discovery import build
import time

def process_inbox():
service = build(‘gmail’, ‘v1’, credentials=authenticate())
emails = fetch_unread_emails()

for email in emails:
category = classify(email[‘body’])
print(f”Processing: {email[‘subject’]} → {category}”)

if category == “urgent”:
# Add label “URGENT” and notify (e.g., send to phone)
add_label(service, email[‘id’], “URGENT”)
send_notification(f”URGENT: {email[‘subject’]}”)

elif category == “reply”:
reply = draft_reply(email[‘body’])
# TODO: Auto-send? Only if you’ve tested thoroughly.
# For now: log and wait.
print(f”[DRAFT] {reply[:100]}…”)

elif category == “archive”:
remove_label(service, email[‘id’], “INBOX”)

time.sleep(0.5) # Respect rate limits (2 req/sec safe)

def add_label(service, msg_id, label_name):
# Create label if not exists, then apply
# (Implementation omitted for brevity—see Gmail API docs)
pass

def remove_label(service, msg_id, label_name):
service.users().messages().modify(
userId=’me’,
id=msg_id,
body={“removeLabelIds”: [label_name]}
).execute()
“`

Run this as a cron job: `*/15 * * * * /usr/bin/python /opt/gmail-ai/inbox_automator.py >> /var/log/gmail-ai.log 2>&1`

—

## Safety & Reliability Checklist

– **Rate limiting**: Gmail blocks IPs that exceed ~2 req/sec. Always throttle.
– **Error handling**: Wrap API calls in `try/except`. Retry transient errors (429, 503).
– **Logging**: Log *everything* (email ID, category, timestamp). Critical for debugging.
– **Human-in-the-loop**: Never auto-delete or auto-send without opt-in. Start with `draft` + `label` only.
– **Model updates**: Re-run classification weekly. Models drift when email patterns change (e.g., new vendors).

> **What doesn’t work?**
> – Parsing HTML-only emails (no plain text fallback)
> – Emails with complex attachments (PDFs, images) without OCR
> – Highly contextual internal comms (e.g., Slack-style threads in email)

—

## Key Takeaways

– ✅ **Use local LLMs**—they’re fast, private, and free for personal use.
– ✅ **Start small**: Classify first, draft later