Tathagata Ghosh

I make tensors smaller and clusters busier.

whoami

  • Tathagata Ghosh
  • M.Tech CDS, IISc Bangalore (2025–27)
  • HPC · CUDA · MPI · multi-GPU tensor SVD
  • Google printed 'Mr. Nasamajh' on my Hash Code certificate. It means clueless. I kept it.

top

  • 90x HYDRA all-pairs shortest paths

    vs my re-implementation of Sao et al.'s 2-D scheme on ego-Facebook: 217.93 s vs 2.42 s, same cluster
  • 0.888 scaling efficiency, 16 GPUs

    thesis: 1024³ tensor, FP32, 14.20x on 16 V100 GPUs, two nodes
  • 3.12x GPU vs 72 CPU threads

    LCS of two 10⁸-char strings: CUDA 0.649 s vs OpenMP 2.028 s
  • 8.7e-15 energy identity error, FP64

    thesis: Theorem 4 / Corollary 6 in FP64 (2.3e-5 in FP32)

cat thesis.md

tSVDm: exact star-M tensor SVD on 16 V100 GPUs

MATRIX Lab, IISc (ACM Gordon Bell Prize-winning lab) · Advisor: Prof. Phani Motamarri · with Argonne National Laboratory · Mar 2026–present

  • C++17/CUDA, MPI + an NCCL ring over NVLink, FP32 and FP64
  • 14.20x on 16 GPUs (efficiency 0.888), 1024³ tensor, FP32
  • energy identity to 8.7e-15 (FP64); 4096³ FP64 ran in 2719 s

ls projects/

  • HYDRA 90x

    All-pairs shortest paths on MPI + OpenMP + CUDA.

    • Density-aware dispatch sends each 32x32 tile to GPU or CPU.
    • Won 8 of 14 runs against 8 re-implemented methods, up to 90x.
    • Matches the sequential reference exactly; all 6 losses reported.
    vs my re-implementation of Sao et al.'s 2-D scheme on ego-Facebook: 217.93 s vs 2.42 s, same cluster.
    • MPI
    • OpenMP
    • CUDA
    DS295 Parallel Programming, grade A+ · Proof: source code (HYDRA)
  • LCS on GPU 3.12x

    Longest common substring of two 10⁸-character strings.

    • No O(NM) table: suffix arrays by prefix doubling on the GPU.
    • CUDA 0.649 s vs best 72-thread OpenMP 2.028 s.
    • 30 trials plus 5 warm-ups; 95% CI ±0.0029 s.
    CUDA prefix doubling 0.649 s vs the best OpenMP configuration, 72 threads, 2.028 s (k-mer index + AVX2).
    • CUDA
    • Thrust
    • AVX2
    DS295 assignment · Proof: source code (LCS on GPU)
  • Face-wise tensor product 18.66x

    Four backends for the product behind t-SVD.

    • Serial, OpenMP, CUDA and hybrid CPU+GPU backends.
    • Up to 18.66x over serial (OpenMP); 3.71x CUDA end to end.
    • Measured across 488 timed runs.
    OpenMP path, 99%-sparse 5000x5000x100 tensor, 64 threads, on a 36-core node with an RTX A5000.
    • C++17 templates
    • OpenMP
    • cuBLAS
    DS285 Tensor Computations, grade A · Proof: source code (Face-wise tensor product)
  • tinyinfer 5.8x

    LLM inference engine from scratch in PyTorch + Triton.

    • KV cache, GQA-aware Triton decode attention, INT8/INT4.
    • CUDA-graph decode 4.8-5.8x faster than Hugging Face generate().
    • Split-KV kernel at 1.68 TB/s, 99% of copy bandwidth.
    Batch-1 decode of Qwen2.5-0.5B and TinyLlama-1.1B on one A100 80GB PCIe, vs Hugging Face generate().
    • PyTorch
    • Triton
    • CUDA graphs
    Personal project, Oct 2026 · Proof: source code (tinyinfer)
  • K-means + SpMV 21.90x

    OpenMP K-means and MPI SpMV on a 760.6 M-nonzero matrix.

    • K-means: 7.19x on 32 threads, same clusters as serial.
    • Showed dynamic scheduling can lose to serial (0.38x).
    • SpMV: 21.90x on 128 processes across 4 nodes.
    Strong scaling of MPI SpMV on the NLPKKT240 matrix, 128 processes across 4 nodes.
    • OpenMP
    • MPI
    • SLURM
    DS295 assignment · Proof: source code (K-means + SpMV)
  • DFT-FE on CLAP

    Routed DFT-FE's 8 GEMM overloads through CLAP's 4.

    • Configure-time switch; bit-identical to upstream when off.
    • CLAP headers isolated behind one shim header.
    • Verification and timing layer on real call shapes.
    • C++
    • CMake
    • cuBLAS
    MATRIX Lab side project · Proof: GitHub profile (DFT-FE on CLAP)
  • Air-B-N-C

    Write characters in thin air; edge devices read them.

    • IMU, camera and depth branches each infer on the edge.
    • A central node fuses them over 62 character classes.
    • Team of 3; I built the IMU firmware and vision classifier.
    • Arduino
    • PyTorch
    • TinyML
    CP330 Edge AI, team of 3 · Proof: source code (Air-B-N-C)
  • SFT, DPO, GRPO

    Post-trained Qwen2.5-0.5B on GSM8K math word problems.

    • SFT, DPO and GRPO objectives written by hand in PyTorch.
    • GRPO: 34.6% to 44.7% greedy on all 1,319 test problems.
    • Traced a 5.5-point DPO drop to likelihood displacement.
    Qwen2.5-0.5B (494M), full GSM8K test set, greedy decoding; GRPO trained 30 min on one A100.
    • PyTorch
    • transformers
    • SLURM
    Rebuilt from DS207, Oct 2026 · Proof: source code (SFT, DPO, GRPO)
  • Wrist IMU HAR

    Activity recognition from a wrist IMU on a Nicla Vision.

    • MicroPython firmware at a drift-corrected 60 Hz.
    • 12 activities, 26,819 windows, 126 features.
    • Chose a 134 KiB, 94.48% tree over a 15 MB forest.
    • MicroPython
    • scikit-learn
    • Optuna
    CP330 Edge AI · Proof: source code (Wrist IMU HAR)
  • VOC vision

    Multi-label classification and segmentation, VOC-style.

    • PCA from scratch: 170 of 3072 components keep 95%.
    • ResNet50 with an FPN-style head: val F1@0.5 = 0.812.
    • ConvNeXt + DeepLabV3+: mIoU 0.604 vs 0.446, p = 7.1e-26.
    • PyTorch
    • timm
    Course project · Proof: source code (VOC vision)
  • Video2PDF

    Lecture videos to deduplicated slide PDFs.

    • Slide changes by pixel-difference ratio, 5-frame stability.
    • Near-duplicates dropped by MSE similarity.
    • Local files or YouTube URLs.
    • Python
    • OpenCV
    Side project · Proof: GitHub profile (Video2PDF)
  • ISDC 2026 site

    TypeScript site for IISc's Durga Puja (ISDC 2026).

    • Built while on the committee, finances included.
    • Also built the CDS farewell 2026 site.
    • TypeScript
    • HTML/CSS
    Community · Proof: GitHub profile (ISDC 2026 site)
  • Trojan GNN

    Golden-reference-free hardware Trojan detection with GNNs.

    • GNNs on data-flow graphs of RTL and gate-level netlists.
    • Clean and Trojan-inserted 16-bit ALU in Verilog.
    • Self-reported: 97% recall at RTL, 84% at gate level.
    • GNNs
    • Verilog
    • TrustHub
    B.Tech · Proof: GitHub profile (Trojan GNN)
  • Logic locking

    Logic locking a 16-bit ALU against Trojan insertion.

    • Cadence Genus/Tempus scripts flag likely insertion sites.
    • XOR/XNOR key gates harden them against SAT attacks.
    • B.Tech final-year project, graded A+.
    • Verilog
    • Cadence
    • Python
    B.Tech final-year project · Proof: GitHub profile (Logic locking)
  • Face login

    Log in with facial key-points instead of a password.

    • Live detection for higher security.
    • Python/OpenCV back end, React + Bootstrap front end.
    • OpenCV
    • React
    B.Tech · Proof: GitHub profile (Face login)
  • IISc innovation archive

    A queryable archive of IISc's innovation history.

    • Five pillars, from startups to eminent faculty.
    • NotebookLM-grounded workflow into one master database.
    • Owned three divisions; interviewed a retired professor.
    • Web scraping
    • NotebookLM
    Institute Innovation Council · Proof: GitHub profile (IISc innovation archive)

history

  • 2020-08B.Tech IT begins, IIEST Shibpur
  • 2023-05Summer research intern, IIEST Shibpur
  • 2024-05B.Tech done, CGPA 9.0
  • 2024-07Software Engineer-1, RMES India
  • 2025-08M.Tech CDS begins, IISc Bangalore
  • 2026-03Thesis: tSVDm, MATRIX Lab
  • 2026-05Data Science Intern, IBM India
  • 2027-06M.Tech ends
  • IIEST Shibpur

    B.Tech IT, 2020–2024

    CGPA 9.0, First Class

    School: ISC 97.75%, ICSE 96.4%

  • IISc Bangalore

    M.Tech CDS, 2025–2027

    CGPA 8.6 · MATRIX Lab

    Thesis: tSVDm, Mar 2026–present

  • Software Engineer-1, RMES India (Rugged Monitoring) 2024–25

    C#/.NET real-time monitoring software

    for high-voltage electrical assets

  • Data Science Intern, IBM India 2026

    Win-likelihood models for four products

    up to 5.7x win-rate lift over baseline

cat skills.txt ranks.txt

  • HPC: CUDA, MPI, NCCL, OpenMP, Nsight / TAU, Roofline
  • AI / ML: PyTorch, XGBoost + SHAP, GNNs, SFT / DPO / GRPO, Optuna, TinyML
  • Systems: C++17, C# / .NET / WPF, Python, TypeScript, Verilog, Linux / Bash
  • AIR 663 GATE DA 2025
  • 1,668,149 pts Hash Code 2022, qualification
  • rank 13,429 Kick Start 2021, Round A

cd ~/offduty

Cycling · Chai · Sleeping · Adda (long, unhurried conversation with friends (Bengali))

  • Hash Code certificate said 'Mr. Nasamajh' (clueless). I kept it.
  • I keep the ledger for the campus cat welfare fund.
  • Mess President for 800+ residents. Lunch is harder than consensus.
  • Asked for 5 timing runs. I ran 30, plus 5 warm-ups.
  • Found two errors in the published pseudocode I was implementing.
  • Built the Durga Puja website, then argued with its invoices.
  • Proved OpenMP dynamic scheduling can lose to serial: 0.38x.
  • Reimplemented eight rival APSP methods. Published all 6 losses.

exit

No email here by design; LinkedIn messages reach me.