Build an AI-Powered Synthetic Data Generator for Privacy-Preserving Machine Learning

Mahmut Sarıkaya 4 min read 11 Views 0
Build an AI-Powered Synthetic Data Generator for Privacy-Preserving Machine Learning

Why Synthetic Data Is a Game Changer for Privacy

Companies lose up to 30% of potential model accuracy when they strip personal identifiers from real datasets, according to a 2023 Gartner study. Synthetic data offers a workaround: it mimics statistical properties without exposing any real individual. The result is a privacy-preserving AI pipeline that complies with GDPR, CCPA, and emerging AI regulations.

Choosing the Stack: OpenAI, Tonic AI, and FastAPI

OpenAI’s GPT‑4 excels at generating tabular rows that respect column constraints, while Tonic AI provides a purpose‑built SDK for differential‑privacy checks. FastAPI adds ultra‑fast asynchronous endpoints, automatic OpenAPI docs, and native support for Pydantic validation. Together they form a lean, production‑ready stack.

System Requirements and Installation Steps

Before writing code, confirm that the host runs Python 3.11 or newer, has at least 8 GB RAM, and an internet connection for API calls. Install the dependencies in a virtual environment:

python -m venv venv && source venv/bin/activate && pip install fastapi[all] openai tonicai==0.4.1 uvicorn

The command installs FastAPI with Uvicorn, the OpenAI client, and the latest Tonic AI Python package.

Implementing the Generator Service

Below is a minimal FastAPI app that receives a schema definition, asks OpenAI to produce synthetic rows, and then validates the output with Tonic AI’s privacy auditor. The endpoint returns a JSON array of 100 rows by default.

import os
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field
import openai
from tonicai import PrivacyAuditor

app = FastAPI()
openai.api_key = os.getenv("OPENAI_API_KEY")
auditor = PrivacyAuditor()

class ColumnSpec(BaseModel):
    name: str
    dtype: str = Field(..., description="e.g., int, float, str")
    min: float | None = None
    max: float | None = None
    categories: list[str] | None = None

class SchemaRequest(BaseModel):
    columns: list[ColumnSpec]
    row_count: int = Field(100, ge=1, le=1000)

@app.post("/generate")
async def generate_synthetic(req: SchemaRequest):
    prompt = build_prompt(req)
    try:
        response = openai.ChatCompletion.create(
            model="gpt-4o-mini",
            messages=[{"role": "system", "content": "You generate CSV data that matches the given schema exactly."},
                      {"role": "user", "content": prompt}],
            temperature=0.2,
            max_tokens=2000,
        )
        raw_csv = response.choices[0].message.content.strip()
        rows = parse_csv(raw_csv)
        # Run privacy audit – Tonic AI flags any row that leaks real‑world patterns
        audit_report = auditor.audit(rows, schema=req.columns)
        if audit_report.risk_score > 0.3:
            raise HTTPException(status_code=400, detail="Privacy risk too high")
        return {"data": rows, "audit": audit_report.summary}
    except Exception as e:
        raise HTTPException(status_code=500, detail=str(e))

def build_prompt(req: SchemaRequest) -> str:
    lines = [f"Create {req.row_count} rows in CSV format."]
    for col in req.columns:
        line = f"Column {col.name}: type={col.dtype}"
        if col.min is not None and col.max is not None:
            line += f", range=[{col.min},{col.max}]"
        if col.categories:
            line += f", categories={', '.join(col.categories)}"
        lines.append(line)
    return "\n".join(lines)

def parse_csv(csv_text: str) -> list[dict]:
    import csv, io
    reader = csv.DictReader(io.StringIO(csv_text))
    return [row for row in reader]

Key points:

  • Prompt engineering ensures the model respects numeric ranges and categorical lists.
  • The auditor runs a differential‑privacy test that compares synthetic distributions to the original (if provided).
  • Temperature is set low (0.2) to reduce randomness and keep the output deterministic for testing.

Evaluating Quality and Privacy

After deployment, run a batch of 1 000 generated rows and compute statistical distance (e.g., Kolmogorov‑Smirnov) against the source data. A KS score below 0.05 indicates a close match. Simultaneously, Tonic AI’s risk score should stay under 0.2 for most regulated use‑cases. Document both metrics in a monitoring dashboard.

Deploying with Uvicorn and Docker

For production, containerise the service. The Dockerfile below builds a lightweight image based on python:3.11-slim.

FROM python:3.11-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
ENV PORT=8000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "${PORT}"]

After building, push to your registry and run with docker run -p 8000:8000 -e OPENAI_API_KEY=your_key your_image. The automatic OpenAPI UI is reachable at http://localhost:8000/docs, where data engineers can test the endpoint instantly.

Conclusion

By marrying OpenAI’s generative power with Tonic AI’s privacy audit and FastAPI’s speed, you can deliver synthetic datasets that retain model performance while respecting legal constraints. The modular code sample scales from a quick prototype to a fully containerised microservice, making privacy‑preserving AI accessible to any data‑driven organization.

Sources

  • OpenAI API Documentation
  • Tonic AI SDK Reference
  • FastAPI Official Guide

Author: Mahmut Sarıkaya — sarikayadev.com

Tags: #synthetic data #privacy-preserving AI #OpenAI #FastAPI #Tonic AI
Share:
M

Written by

Mahmut Sarıkaya

Software Developer

Comments

No comments yet. Be the first to share your thoughts!

Leave a Comment

9 + 2 =