Software Engineer · GenAI · Full-Stack
United States · prakritimksharma@gmail.com · +1 (312) 479-6193 · github: github.com/prakriti31 · linkedin: linkedin.com/in/prakritisharma31/
Master of Science in Computer Science, CGPA: 3.5/4.0
Bachelor of Technology in Computer Science Engineering, CGPA: 8.56/10.0
Caregiver marketplaces lean on opaque, one-time background checks, leaving families under-protected when vetting childcare and eldercare providers. Wardwell is a two-sided marketplace built around a transparent, continuous trust layer instead — plain-language tiered verification badges, a visible "continuously monitored since [date]" indicator backed by Checkr's continuous-check product, and structured reference checks, all instead of the industry's one-time badge. Booking and messaging stay on-platform until both sides opt in, and a trust-ops admin dashboard gives staff a resource-based, auth-gated view into every verification and incident. Runs on sandboxed Checkr, Stripe Identity, and Stripe Connect integrations end to end.
Postgres, MySQL, MongoDB, and DuckDB offer no native point-in-time history, forcing teams to hand-build audit tables or run CDC pipelines just to answer "what did this row look like last Tuesday." Stratum is a single-file embedded database that auto-versions every write with SQLite's ergonomics and git's memory, and exposes time-travel SQL verbs — AS OF, DIFF, HISTORY, BLAME, ROLLBACK — through a CLI, REPL, and web UI, with zero config required to turn it on.
Git's diff and merge machinery treats every file as lines of text, which breaks down hard for structured formats — two engineers editing unrelated cells in the same Jupyter notebook still conflict, because array insertions shift every line after them and JSON has no concept of "this cell is a unit." semdiff is a general, pluggable, conflict-aware structural merge engine: a generic tree representation, an identity/content-hash/GumTree-style matcher, and format adapters (tree-sitter for JSON and Terraform HCL, nbformat for notebooks, ruamel.yaml for Kubernetes manifests) built once so a new format is a parser plus semantic rules, not a new project. Non-overlapping edits merge cleanly where git falsely conflicts, real conflicts come back as valid, still-parseable files with inline markers instead of raw <<<<<<< text, and it installs as a one-command git merge-driver across all four formats.
l2_squared, dot_product(real), and cosine_similarity(real) silently failed to register in default (non-FAISS) Velox builds, breaking real-column queries with "Scalar function name not registered" — while calls against literal arrays appeared to work, since the Presto coordinator constant-folds literal expressions before they ever reach the native worker. Authored pure C++ fallback implementations for all four function/type overloads, registered them unconditionally regardless of VELOX_ENABLE_FAISS, and un-gated the matching unit tests so they run in default builds instead of only under FAISS.
The iterator-debugging proxy machinery in <xmemory> — _Container_proxy, _Container_base12, _Iterator_base12 — governs how MSVC's debug iterators stay valid across container mutation, but carried no comments despite genuinely non-obvious invariants: an intrusive linked list of live iterators, mutual container/proxy pointers, and a specific locking rule. The explanation had lived only in a maintainer's issue-thread reply for four years. Verified those invariants still matched the current code exactly, then wrote them up as comments next to the code they describe — a comment-only PR with no behavior change.
A RAG system that surfaces not just what a policy says, but why it exists. Sludge-RAG mines the Slack threads and email chains around a document, runs a multi-turn Claude pipeline to extract the organizational rationale behind each decision, and attaches it as metadata to every indexed chunk — so a query like "why is retention 90 days" comes back with the GDPR and audit context, not just a dry PDF excerpt. A dual-pane "Vibe-Check" playground puts plain vector search side by side with rationale-augmented search so the difference is visible, not just claimed.
A SaaS dashboard for developers running multiple AI coding agents — Claude Code, Copilot, Cursor, Codex — in one place instead of juggling terminals and logs. Tracks sessions per agent with token counts, cost, duration, and status, streams a terminal-style activity log, and pushes Slack and email alerts when a session finishes or fails. Auth is JWT plus Google OAuth with per-user @alias handles for team visibility, rate-limited and request-logged on the backend.
A multi-tenant expense tracking and splitting SaaS with AI categorization, a conversational assistant, and real subscription billing — not a toy CRUD app. Seven FastAPI microservices talk over a RabbitMQ event bus with dead-letter queues and per-consumer dedup, Gemini categorizes expenses and answers questions, and Stripe Checkout handles billing with signature-verified, idempotent webhooks. Runs on Kubernetes (EKS) via Terraform-provisioned infra, autoscales on CPU and queue depth, and ships through a GitHub Actions pipeline with coverage gates, Trivy scanning, and Prometheus/Grafana/Loki observability.
An Instagram analytics pipeline that scrapes profiles via Apify, queues the work through Kafka, and pushes results to a React dashboard over WebSockets in real time. FastAPI handles the API layer, a Python consumer service does the actual scraping and writes to MongoDB or DynamoDB depending on deployment, and rate limiting plus per-query performance logging keep the scraping layer honest at scale.
A between-visit clinical intelligence platform for hypertension management that gives clinicians an 8-minute pre-visit briefing instead of a blank chart. Ingests patient data via FHIR R4, then runs a three-layer nightly pipeline — a deterministic rule engine (gap, therapeutic inertia, adherence-BP correlation, deterioration, and variability detectors) feeds a weighted risk score, which feeds a Claude-generated 3-sentence clinical narrative validated against faithfulness and safety guardrails before it's ever shown. Also runs a deterministic drug-interaction checker, a tool-using clinical chatbot, and a patient-facing PWA for home BP submission — all driven by a 30-second poll worker, no Celery, no Redis queue.
A scoped AI chat assistant for PartSelect's refrigerator and dishwasher catalog — finds parts by symptom or model number, verifies compatibility against a specific model number (never guesses), and walks customers through installation and troubleshooting with interactive, localStorage-persisted checklists. Built on Claude tool-use for part search, compatibility checks, and order support, with rich inline product cards and hard scope enforcement so it refuses anything outside the parts catalog.
A shadow removal algorithm that detects shadows and lifts them by matching each shadow region to its most similar non-shadow counterpart, then transferring brightness while preserving texture. Detection combines gradient, texture (Texton), and spatial-distance features from mean-shift segmentation with Y-channel and HSI-space brightness cues clustered via k-means. Removal uses HSV histogram matching between each shadow region and its best-matched non-shadow region, followed by boundary smoothing across the seams. Built and evaluated in MATLAB.
Languages: Java, Python, JavaScript, TypeScript, Go, C#, C++, HTML5, CSS3
Frontend & Backend: React, Node.js, Next.js, Angular, Vue.js, Express.js, Spring Boot, .NET Core, Django, FastAPI
Cloud & DevOps: AWS, GCP, Docker, Kubernetes, Jenkins, Kafka, CI/CD, Infrastructure Automation
Databases & Data: MySQL, PostgreSQL, MongoDB, DynamoDB, SQLite, Elasticsearch, GraphQL, Vector Databases, Embedding Search
AI / GenAI: LangChain, OpenAI API, Gemini, Hugging Face, Pandas, Scikit-learn, RAG, Prompt Engineering, NLP, Vector Search, Generative AI, AI Agents, LLM Integration, Embeddings
Testing & QA: JMeter, Appium, BrowserStack
Integrations & Tooling: Stripe, Clerk, Checkr, Prisma ORM, CMake, GTest, Vite