Skip to content
View Mullassery's full-sized avatar

Block or report Mullassery

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Mullassery/README.md

πŸ‘‹ Hi, I'm Georgi Mullassery

MarTech Solutions Architect | Customer Data Platforms β€’ Data Engineering β€’ AI Engineering | Bengaluru, India

Email LinkedIn Product Hunt GitHub

πŸš€ About Me

Solutions Architect with more than a decade of work experience across IBM, Wipro, IPG Mediabrands, and others, designing and delivering enterprise-scale customer data, marketing technology, analytics, and AI-enabled solutions. Expertise in Customer Data Platforms (Adobe RT-CDP, Braze, mParticle, Segment CDP), event-driven architectures, cloud data platforms, customer journey orchestration, identity resolution, real-time activation, and data engineering on Google Cloud and Azure.

I build open-source Python libraries that solve the practical problems around running a Customer Data Platform in production β€” tag/analytics implementation, audience intelligence, reverse ETL, and the data-pipeline reliability that CDPs depend on. Outside of MarTech, I also build systems-level software for fun (an AI-native OS kernel written in Rust β€” the SHER OS family).

  • πŸ”­ Current focus: Customer Data Platforms, marketing data pipelines, and the Python tooling that supports them
  • πŸ€– Exploring: Agentic AI for MarTech workflows, RAG architectures, MCP-based agent tooling
  • πŸŽ“ Background: MBA, Jansons School of Business Β· B.Com, Mar Ivanios College

πŸ”§ Technologies & Tools

MarTech & Customer Data Platforms Adobe RT-CDP, Braze, mParticle, Segment CDP, customer journey orchestration, identity resolution, real-time activation, Hightouch / Reverse ETL

Data Warehousing & Engineering BigQuery, Snowflake, Databricks (PySpark), Microsoft Fabric, Azure Data Factory, Airbyte / Fivetran, dbt, Apache Airflow, Apache Kafka, Apache Flink (PyFlink), InfluxDB, Telegraf, Datadog

Cloud & Data Platforms GCP, Azure, AWS Β· Google Cloud Storage, Azure Blob Storage / ADLS, Amazon S3, MinIO

AI β€” RAG & Application Development LangChain, LangGraph, LlamaIndex, Chainlit / OpenWebUI, LangSmith, DSPy, BM25 + vector similarity + reranking, Context Engineering, Ragas + DeepEval, Pinecone, Weaviate, ChromaDB, pgvector, Qdrant, Milvus

AI β€” Agent Automation Temporal, Langflow, n8n + Node.js, FastAPI, Firecrawl, Tavily, SerpAPI, Twilio, Claude Agent SDK, OpenAI Agent SDK, Google Agent SDK, Ollama, Scheduled & Background Agents

AI β€” Autonomous Code Generation Claude Code, Cursor, Replit, Codex, Hermes Agent, OpenCode, Loveable, OpenRouter

MCP & Model Platforms Model Context Protocol, Hugging Face, AWS Bedrock, Azure AI Foundry, GCP Vertex AI (Model Garden)

Infrastructure as Code & Containers Terraform, Docker, Docker Compose, Azure Kubernetes Service (AKS), NGINX Ingress

Identity & Security Entra ID / Azure AD, Azure Key Vault, Azure AI Search, incoming/outgoing guardrails

IoT, Messaging & Robotics Raspberry Pi 5, ESP32, MQTT (Eclipse / RabbitMQ), Node-RED, AWS IoT Core, AWS IoT SiteWise, CoAP, Edge Processing, Industrial IoT

🌟 Featured Projects β€” MarTech & CDP Tooling

🏷️ PyTagManager β€” AI-native analytics implementation

Crawls a site, builds a semantic DOM graph, and exports ready-to-use tracking configs to GTM, GA4, Segment, Snowplow, Tealium, RudderStack, and Adobe Tags β€” no manual CSS-selector hunting.

πŸ‘₯ ClusterAudienceKit β€” Enterprise audience intelligence

RFM analysis, 6 clustering algorithms, CLV, churn detection, and lookalike modeling. Processes 1M+ customers in under a second.

πŸ” PyReverseETL β€” Quality-validated data activation

Moves data to where it's needed with a full audit trail, lineage tracking, and compliance records β€” the activation counterpart to CDP ingestion.

βœ… PyAirflowTester & StatGuardian β€” Data pipeline reliability

Dependency intelligence and declarative data-quality validation (schema validation, drift detection, anomaly detection) for the Airflow/dbt pipelines that feed customer data platforms.

Also building outside the MarTech day job: SHER OS, an AI-native operating system kernel in Rust, and TinyBridge, a macOS-native Linux VM runtime β€” see the full index below.

πŸ“‚ All Repositories

MarTech & customer data
Repo Description
PyTagManager AI-native analytics implementation β€” semantic DOM graph to GTM/GA4/Segment/etc.
ClusterAudienceKit Enterprise audience intelligence β€” RFM, clustering, CLV, churn detection
PyReverseETL Quality-validated reverse ETL / data activation with lineage tracking
Data pipeline reliability & engineering
Repo Description
PyAirflowTester Airflow & dbt reliability and quality-assurance platform
StatGuardian Declarative data quality framework (Rust) β€” 13x faster than pandera
PyBeamGuard Apache Beam & Dataflow pipeline analysis and cost forecasting
PyDependencyCheck Dependency intelligence for Python β€” supply chain integrity
PyStreamXL Stream large Excel files with constant memory β€” 46x faster than openpyxl
PySynthData Synthetic data generation for ML training
PyWeatherEnriched Weather data enrichment for ML pipelines
PyNetworkIntel Network discovery, topology mapping, anomaly detection
LLM & AI tooling
Repo Description
PyTokenCalc Token counting & cost estimation across 20+ LLM providers
PyInferenceManager Multi-provider LLM inference executor β€” 11+ providers, batching, load testing
PyStreamMCP Query planning & context discovery for AI agents β€” 60–75% token reduction
PyVectorHound Diagnostic engine for RAG retrieval failures
PyStreamPDF Selective PDF extraction to cut RAG costs 50–70%
OpenAnchor Token intelligence middleware for multi-provider LLM usage
PyAPICheck Transparent API security discovery and policy generation
Systems & runtimes (outside the MarTech day job)
Repo Description
TinyBridge macOS-native Linux VM runtime on Apple's Virtualization.framework
PrismNote Jupyter-compatible data-science notebook with built-in AI
homebrew-tinybridge Homebrew tap for TinyBridge
homebrew-prismnote Homebrew tap for PrismNote
SHER OS platform (AI-native operating system, Rust)
Repo Description
SHER-KERNEL AI-native OS kernel β€” zero-trust security, isolated driver runtime, Linux kernel interface
SHER-Graphics Software GPU simulation + real Vulkan/MoltenVK backend
SHER-Display Display server, compositor, and window manager
SHER-INPUT Input subsystem β€” canonical event stream for Display and Aurora
SHER-Aurora GNOME-style design system: tokens, typography, motion, accessibility
SHER-Process-Explorer Evidence-based Linux process explorer β€” /proc telemetry, eBPF sampling, daemon + CLI + desktop UI
Robotics & simulation
Repo Description
PyRoboFrames High-performance ML dataloader for robotics (LeRobot format, Apple Silicon)
PyRoboReplay Robotics perception and replay engine β€” sensor fusion, trajectory analysis
PyRoboSimulator World simulator for autonomous systems β€” 100K+ agents, REST API
PyRoboVision Perception stack for autonomous robots & vehicles
PyTerrainMap Unified terrain intelligence for multi-robot fleets

πŸ“« Reach Me

Thanks for stopping by β€” always happy to talk CDPs, marketing data pipelines, or the Python tooling that makes them reliable.

Pinned Loading

  1. ClusterAudienceKit ClusterAudienceKit Public

    Enterprise audience intelligence at scale. RFM analysis, 6 clustering algorithms, CLV, churn detection, lookalikes, neural networks. Process 1M+ customers in <1s.

    Python

  2. PyAirflowTester PyAirflowTester Public

    Enterprise Airflow & dbt Reliability Platform. Complete dependency intelligence and quality assurance system for modern data platforms.

    Python

  3. PyDependencyCheck PyDependencyCheck Public

    Production-grade dependency intelligence for Python. Why dependencies exist, who added them, if they're used, if they're safe, what changed. Supply chain integrity.

    Python

  4. PyReverseETL PyReverseETL Public

    Move your data automatically to where it's needed. Quality-validated data activation with full audit trail, lineage tracking, and compliance records.

    Python

  5. PyTagManager PyTagManager Public

    AI-native analytics implementation platform: crawl a site, build a semantic DOM graph, and export ready-to-use tracking configs to GTM, GA4, Segment, Snowplow, Tealium, RudderStack, and Adobe Tags …

    Python

  6. StatGuardian StatGuardian Public

    Fast, declarative data quality framework (Rust). Schema validation, drift detection, anomaly detection. 13x faster than pandera. Pandas/Polars/DuckDB support.

    Python