Experi v2, Production A/B Testing on Databricks
Databricks CUPED Bayesian Inference MLflow Delta Lake Gradio
Rebuilt the backend of Experi on Databricks. Raw events land in Delta Lake, with CUPED variance reduction, Bayesian inference, and full experiment logging via MLflow 3.0. Every result is persisted to Unity Catalog with full lineage, designed to be reproducible and auditable.
Live at theshlok13-experi.hf.space < View on GitHub <
Referral Copilot, Databricks Data + AI Summit Hackathon
Databricks Apps Relevance Scoring Healthcare Python
June 2026
Built an AI agent workflow for healthcare referral-ranking at a 480-hacker, OpenAI-sponsored hackathon (Databricks Data + AI Summit, Apps & Agents for Good track): a Databricks App matching patients to healthcare facilities by care need and location, ranked by an embeddings-based relevance-scoring model across 10,000+ facility records. Followed Git version control, CI/CD, and LLMOps-style monitoring practices to ship a working demo end-to-end within the event window.
Top 7 finish
View on GitHub <
Experi v1, Browser-Based Experiment Design Tool
Experimentation CUPED Bayesian Sequential Testing Vanilla JS
Built and deployed a free experiment design tool for startup PMs. Input your CVR and traffic to get sample size, runtime, CUPED variance reduction, Bayesian probability, sequential testing boundaries, live experiment analyzer, and an experiment risk score out of 100 that grades your setup across 5 dimensions before you launch.
43.8M earned media
Risk score: 97/100 on default config
Live at tryexperi.netlify.app <
Product Experiment and A/B Test Analysis
Experimentation Hypothesis Testing Power Analysis Guardrail Metrics Python
Mar – Apr 2026
Designed and analyzed a controlled A/B experiment identifying a statistically significant 35% engagement uplift using hypothesis testing, p-values, and confidence intervals. Defined guardrail KPIs to monitor for unintended negative effects, validated the result with a power analysis for sample sizing, and documented a clear rollout recommendation.
35% engagement uplift
View on GitHub <
Causal Inference in Product Analytics
PSM Diff-in-Diff Uplift Modeling Python
March 2026
Cross-validated a true treatment effect estimate across three methods (propensity score matching, difference-in-differences, T-learner uplift modeling) on a non-randomized feature rollout across 100K users. Root cause analysis uncovered that the naive estimate had overstated the causal effect by 85.7%, with the uplift model built from scratch using scikit-learn. Checked that all three causal methods converged on a consistent estimate before trusting the result.
85.7% bias removed via PSM
View on GitHub <