From building the first Kashmiri LLM to shipping full-stack products with real users — I solve hard problems at the intersection of AI and software engineering.
CS Engineering student · AI Engineer · Full-Stack Builder
Not side projects. Real systems built to solve real problems at scale.
7 million+ people speak Kashmiri — yet no large language model existed for the language. KoshurAI is the first Kashmiri LLM, built from scratch with a 150K+ human-verified translation pair dataset and fine-tuned on 2M+ tokens. Presented at AI Impact Summit 2026.
Kashmiri is a low-resource language spoken by 7M+ people with near-zero representation in modern AI systems. No translation datasets, no NLP tools, no LLM — leaving an entire linguistic community excluded from the AI revolution.
Kashmiri text frequently omits diacritics, creating lexical ambiguity that breaks every downstream NLP task. Koshur Diacritizer is a ByT5-small seq2seq model that restores them — published on arXiv, released on Hugging Face, and covered by ETV Bharat.
The same undiacritized Kashmiri surface form can map to several distinct words. Tokenizers, translation systems, search, and TTS all degrade without diacritics — and no neural solution existed for Kashmiri.
Five architectural generations across ~11 experimental iterations on a ~673K-sample training ecosystem — the journey from a CNN baseline to foundation-model OCR, and the discovery that re-pointed the whole program.
Each generation eliminated an architecture family: CNNs lacked sequence understanding, CTC plateaued on ligatures, PARSeq failed to generalize to Nastaliq, and TrOCR hit tokenizer walls. v5 (Arabic-TrOCR fine-tune) works best — and exposed decoder sequence collapse: the encoder learns Kashmiri glyphs; the decoder truncates long sequences. The bottleneck is no longer vision, it's decoding.
The first large-scale synthetic OCR dataset for Kashmiri Nastaliq — 613,078 image-text pairs generated with the SynthOCR-Gen framework, published on arXiv. The data-side intervention that powers the KoshurOCR research program.
OCR for low-resource scripts is gated by data, not architecture. Manual annotation of Nastaliq doesn't scale — contextual ligatures make character-level labeling expensive and error-prone. Synthetic generation is the only path to scale.
Can frontier LLMs read sensor data and reason correctly about physical danger? A 68-scenario benchmark evaluating 5 frontier models across toxic gas, thermal, mechanical, intrusion, and deliberately ambiguous edge cases. Paper under review.
LLMs are being wired into industrial monitoring, smart buildings, and robotic safety systems — but no public benchmark tests whether they can turn multi-sensor readings into correct hazard assessments and appropriate actions. Existing benchmarks test academic reasoning, not physical-world judgment.
A scalable real-time media sharing platform built around instant media uploads, room-based sharing, and live collaboration. Architected with Supabase for real-time sync and Cloudflare R2 for globally distributed media delivery.
Existing media sharing tools are bloated, slow, and not built for real-time group contexts. Groopik is purpose-built for instant, room-based collaborative media sharing with zero friction.
A voice-first AI assistant that helps users craft precision prompts through iterative clarification. PromptForge guides users from vague intent to optimized, structured prompts — reducing LLM output variance and improving task completion rates.
Most users get poor results from LLMs not because the model is bad — but because their prompts are imprecise. PromptForge solves prompt quality at the input layer through voice and structured clarification.
3 arXiv preprints and an ongoing multi-generation OCR research program — spanning NLP, computer vision, and LLM evaluation.
ByT5-small model for restoring diacritics in Kashmiri text — DERm 0.0212, WER 0.2159, 77.5% expert accuracy. Includes the first public dataset of 23.7K aligned diacritized Kashmiri sentence pairs. Covered by ETV Bharat.
View on arXiv ↗ Model on HuggingFace ↗ Press · ETV Bharat ↗613,078 image-text pairs for Kashmiri OCR, generated via SynthOCR-Gen from the KS-PRET-5M corpus. First large-scale synthetic OCR dataset for the Kashmiri script, covering 25+ augmentation strategies across multiple fonts and granularities.
View on arXiv ↗ View on GitHub ↗68-scenario benchmark evaluating 5 frontier models on physical hazard reasoning across multi-sensor contexts.
View on GitHub ↗Five architectural generations (CNN → CRNN+CTC → PARSeq → TrOCR → Arabic-TrOCR) across ~11 experimental iterations on a ~673K-sample training ecosystem. Key finding: decoder sequence collapse under heterogeneous supervision — the bottleneck is no longer visual representation learning but robust long-sequence autoregressive decoding. Released Koshur-OCR-Synth, a solo-built 571,740-pair clean synthetic OCR dataset.
View on GitHub ↗ Dataset on HuggingFace ↗AI summits, national podiums, state championships. Each achievement is proof, not decoration.
From LLM fine-tuning to full-stack product engineering.
Active engineering threads right now.
Remediating decoder sequence collapse in Nastaliq OCR — length-aware decoding objectives and short→long curriculum learning over the 673K-sample ecosystem.
ActiveSpeech + RAG over the Kashmiri corpus — the next pillar of the Kashmiri AI stack, building on KoshurAI, the Diacritizer, and OCR.
🔬 PrototypingContributing to AI/agent projects via GirlScript Summer of Code 2026, while iterating on Groopik and PromptForge with real user feedback.
ShippingI'm a builder from Kashmir who enjoys solving difficult problems through AI, full-stack software, and product engineering. I don't believe in building things for the sake of building — every project I take on has a clear problem statement and a measurable outcome.
From creating the first Kashmiri LLM to shipping full-stack products with real users and publishing AI research, I focus on work that creates measurable impact for real people.
I'm equally at home training a model on CUDA clusters, designing a PCB for a combat robot, or architecting a scalable backend. That cross-domain fluency is what I bring to every team.
The discipline that wins Wushu championships is the same discipline that ships code at 2am.
J&K State Level
National IIT Podiums
AI Summits & Events
15 AI Projects Shown
Archery · Football · Swim
Press, events, and stages that created credibility beyond the code.
Featured by Brut India as university representative at AI Impact Summit 2026 — presenting 15 AI projects. National digital coverage reaching millions.
Selected as university representative to present 15 AI projects at national AI Impact Summit 2026.
Met Chief Minister of Punjab at the Startup Punjab Conclave — rooms most students never enter.
Top-ranked competitor at IIT Mandi's Xpecto'25 robotics championship among 1000+ participants.
National podium finish at IIT Jodhpur's Prometeo'25 robotics combat competition.
"Three Kashmir Engineers Develop AI Tool to Preserve Native Language in Digital Age" — national coverage of our Kashmiri language AI work (KoshurAI · Koshur Diacritizer).
Coordinated robotics outreach with Indian Army through RISC club as Outreach Head, LPU.
30+ certificates from globally recognized platforms. A selection of the most relevant.
Let's talk.