rohit@rajan:~$
New Delhi, India · open to new roles

Rohit Rajan

Tech Lead · Forward-Deployed Engineer · Product-minded
Senior Engineer & Technology Manager @ Sparklin Innovations

  • I ship production systems in TypeScript, Node.js and PostgreSQL, and own them from design through on-call.
  • I lead a 14-person engineering team and still write production code every day. On the side: RAG pipelines, embedding services, vector search, and an open-source SRE agent framework I helped maintain.
  • Looking for tech lead and forward-deployed engineer roles where I shape the product as much as the code.
8+ yrsshipping production systems
14engineers led, still coding daily
500kmonthly users on products I architect
60–70%compute cost cut, AWS to hybrid
99.95%uptime on AWS workloads

project walkthroughs

How the work actually got built

The problem, the architecture, the calls I made, and what broke along the way. Open any project for the full walkthrough.

click or press Enter to expand

Problem

Our consumer products needed a personalized feed that stayed fast under real production traffic.

Hard decisions

  1. Fan-out, asyncFeeds are built by an asynchronous fan-out pipeline instead of being assembled on each request.
  2. Cache processed feeds in RedisServing pre-processed feeds from Redis is what holds response times under 50 ms in production.
  3. Fix the read path, not just the symptomAfter the outage: paginated queries and database-level indexes instead of unbounded collection reads, then a sweep of the codebase for the same pattern.

What broke

The feed service started timing out, but only for our largest accounts. I traced it to unbounded collection queries and N+1 lookups, redesigned the read path with paginated queries and database-level indexes, and brought tail latency back under SLA the same day.

Then I swept the codebase for the same anti-pattern and found two more services with it before they turned into incidents.

Architecture

green dashes = hot read path

Outcomes

<50 msresponses under production load
Same daytail latency back under SLA
2more services fixed before they became incidents

Problem

A legacy Node.js monolith sat under three products and had to be decomposed to support a wider digital transformation, without taking any of them offline.

Hard decisions

  1. Split by domain ownershipServices were carved out along domain lines, each owned by a team.
  2. Tenant-by-tenant cutover with reconciliation gatesNo big-bang switch. Each tenant moved only after reconciliation checks passed.
  3. Legacy stays the source of truthThe legacy path stayed live as the source of truth until the new path passed two full business cycles. I wrote the migration playbook the team followed.

Architecture

migration flow, one tenant at a time

Outcomes

3products migrated
4 monthsend to end
2 cyclesfull business cycles passed before trusting the new path
Playbookwritten and reused for workflow systems

Problem

Production compute ran entirely on AWS (EC2, S3). I own the deployment, monitoring and scaling story for a consumer product with real user traffic, so cost and headroom were mine to fix.

Hard decisions

  1. Hybrid, not a full exitCompute moved to Hetzner bare metal while AWS stayed part of the setup.
  2. Nginx in front of everythingA single Nginx layer fronts both sides of the hybrid.

Architecture

after the migration

Outcomes

60–70%lower compute cost
Moreheadroom for every service

Problem

A RAG pipeline over an internal corpus. Handing back whole documents isn't precise enough, so it returns the specific passage with its surrounding context, and it has to respect time.

Retrieval quality is graded by evaluation scripts against a held-out question set, not by feel.

Hard decisions

  1. Embeddings as a separate FastAPI async workerThe Gemini 768-dimensional embedding service batches and retries independently from ingest, so upstream timeouts and rate limits never block uploads.
  2. FAISS to QdrantMoved production vector search to Qdrant so metadata filtering happens inside HNSW traversal instead of as post-filtering.
  3. Measure before trustingAn evaluation harness scores retrieval strategies against a held-out question set.

What broke

Users with large libraries were getting near-empty result pages. On FAISS, metadata filters ran after the vector search, so the top results were often thrown away by the filter.

Moving to Qdrant put the filter inside HNSW traversal, which fixed it.

Architecture

ingest (top) · query (bottom) · eval grades retrieval

Outcomes

+35%retrieval accuracy
768-dGemini embeddings, batched + retried
Fixednear-empty result pages for large libraries
0uploads blocked by upstream timeouts or rate limits

Problem

Tracer-Cloud/opensre is a LangGraph-based framework for building AI SRE agents. Agents investigating incidents need to read production logs, and teams running self-hosted Elasticsearch had no log source for it.

This integration lets downstream agents query production logs during incident investigations.

Hard decisions

  1. Ship it as a LangChain toolThe integration is wrapped as a LangChain tool so agents in the framework can call it directly.
  2. A small httpx clientFive methods against Elasticsearch, kept narrow and fully covered by 31 unit tests.
  3. Tests before featuresA separate test-only PR took the auth and AWS modules from 0% to 100% coverage with 34 unit tests, making future contributions safer to review.

Architecture

where the integration sits in the agent

Outcomes

2merged PRs
0→100%coverage on auth.py + aws.py
Paidmaintainer role, 5–8 hrs a week (past)

Problem

Commitments get made in WhatsApp chats and then lost. This captures them where they happen and puts a countdown on your Mac's menu bar.

TickDown itself is an open-source macOS menu bar app that counts down your day: seconds, minutes and hours left today, this week, this month and this year.

Hard decisions

  1. Read the WhatsApp Web DOMThe extension captures commitments straight from the page you're already using.
  2. One shared Convex backendThe extension and the macOS app sync through the same Convex backend.
  3. Rust FFI into SwiftA Rust FFI bridge connects into the native Swift/SwiftUI app.

Architecture

cross-device sync path

Outcomes

DMGTickDown shipped via GitHub releases
OSSTickDown is open source

Problem

Aviation enthusiasts and flight-sim pilots want to see aircraft and weather on the same map. Flightradar24 keeps weather behind a subscription.

The extension adds four live layers you toggle from a popup, with adjustable opacity and data timestamps.

Hard decisions

  1. Zero dependencies, no build stepPlain JavaScript on Manifest V3, about 1,300 lines in total. Load unpacked and it runs.
  2. Isolated contexts with a narrow bridgepage.js draws on the map, a bridge relays messages, and the service worker does the network work. Page requests are limited to named data types, so the page can never make the extension fetch an arbitrary URL.
  3. Privacy firstOnly the storage permission, talks only to MSN's weather servers, collects nothing.
  4. Cache in the service workerThe worker fetches, caches and decodes weather data, and caches map metadata for 3 minutes.

What can break

The data source is undocumented. MSN Weather's public map feed can change or break without notice, and runs 10 to 20 minutes behind real time.

Forecast frames use a proprietary binary format, so only current precipitation is shown. The overlay also depends on Flightradar24 continuing to use Google Maps. Not for real-world navigation or flight planning.

Architecture

content script · bridge · service worker · data feed

Outcomes

4layers: radar, lightning, storm cells, satellite
~1,300lines, zero dependencies
1permission requested (storage)
AnyChromium browser: Chrome, Edge, Brave, Arc and more

experience

Where I've worked

From C# and PHP to leading a team and owning architecture for consumer products at scale.

leading

Lead and still ship

I run a 14-person team across backend and frontend. I mentor junior, mid-level and senior engineers through weekly 1:1s and code-review pairing, and run end-to-end hiring for engineering roles.

product

Product is part of the job

I partner closely with product, design and QA, and turn ambiguous product asks into engineering briefs with concrete options and latency targets. I own features from ideation and prototyping through production and on-call.

  1. Dec 2019 – Present · New Delhi

    Senior Engineer & Technology Manager · Sparklin Innovations Pvt. Ltd.

    Lead a 14-person engineering team that delivers 40% faster, own the architecture for consumer products serving 500k monthly users, and still write production code daily.

    • Own PostgreSQL schema design, indexing and query optimization, and hand-write complex SQL when ORM-generated queries are too slow.
    • Designed structured logging and telemetry around critical paths (latency percentiles, failure reasons, traces) so regressions surface as a data problem instead of a user-complaint problem.
    • Ran production workloads on AWS with Docker, Nginx and GitHub Actions at 99.95% uptime.
    • Internal tools built with product, design and QA cut feature cycle time 30%.
    • Reporting pipeline automation cut manual effort 65% across 4 teams; finance workflow automation saves 120 hours a month.
    • Checked our deployed dependency graph against the axios npm supply-chain incident, confirmed none of our servers were compromised, and added a lightweight dependency review step to releases.
    • Run code review for backend PRs and manage releases to production with unit and integration tests and CI/CD on GitHub Actions.
  2. Past · alongside full-time role · ~5–8 hrs/week

    Paid Maintainer · Tracer-Cloud/opensre

    PR review, issue triage and feature work on a LangGraph-based SRE agent framework. Shipped the Elasticsearch log source and 65 unit tests across two PRs.

  3. Jul 2019 – Oct 2019

    Full Stack Web Developer · Kantag Solutions Pvt. Ltd.

    Built custom e-commerce platforms in PHP and Laravel with end-to-end payment gateway integration and inventory management.

    • Wrote hand-tuned SQL for sales reporting where ORM-generated queries were too slow, keeping client dashboards responsive at peak query times.
  4. Oct 2017 – Jun 2019

    Web Developer (PHP) · TEMZ

    Built core modules for a custom CRM and designed REST APIs consumed in production by partner mobile apps and external integrations.

  5. Feb 2017 – Jun 2017

    Junior C# Developer · Caresoft Inc.

    Maintained Windows Forms desktop applications, handled routine database maintenance and wrote stored procedures for reporting.

  6. 2013 – 2017

    B.E., Computer Science · SAM College of Engineering & Technology

skills

What I work with

Grouped by where it shows up in my work. No progress bars.

languages

TypeScriptJavaScriptPythonPHPC#

backend

Node.jsREST APIsFastAPIDjangoasyncioLaravelMicroservicesService-oriented architectureMonolith decompositionAsync workersBatched retry / backoff

data

PostgreSQLComplex SQLRedis / ValkeyMongoDBElasticsearchQdrantFAISS

cloud_infra

AWS EC2AWS S3ALBCloudFrontDockerNginxLinuxGitHub ActionsHetznerBare-metal deployment

ai_llm

LangChainLangGraphRAG + eval harnessesGemini embeddingsOpenAI CLIPVector search internalsLLM observabilityMCP tool wrappersOpenAI APIAnthropic API

frontend_and_clients

ReactResponsive UIChrome extensions (MV3)Cross-device state sync

reliability

Production profilingMonitoringStructured loggingLatency percentilesDebugging under loadApplication securitySupply-chain review

leadership

1:1sMentoringHiringProduct partnershipEngineering briefsTechnical writing

practice

TDDUnit + integration testsRelease managementCode review cultureBlameless postmortemsDesign docsRunbooksDependency reviewSecure developmentRemote / async

daily_tools

Claude CodeCursorVS Code

contact

Hiring a tech lead or forward-deployed engineer?

Email or call. I'm in New Delhi (IST, UTC+5:30).

Download resume (PDF)