Muhammad Shah
Dubai --:-- GMT+4

AI Software Engineer · HumAI, Dubai · 2025 —

I build AI systems that keep working after the demo.

I design and ship the services behind a conversational CRM assistant, voice agents, an enterprise knowledge base and a customer context graph — LangGraph and DSPy orchestration through to continuous release on Kubernetes — and the benchmarking, regression tests and monitoring that keep language-model systems dependable once they are live.

01

Now — building Huscribe at HumAI

Jan 2025 – present · Dubai

Principal contributor to Huscribe, HumAI’s AI platform: five production services in Python and FastAPI, orchestrated with LangGraph and DSPy, released continuously to Google Kubernetes Engine. The work below spans all of them.

Huscribe · talk to your CRM

The reasoning layer behind a conversational CRM

Automatic field resolution and record normalisation across Salesforce, HubSpot and Zoho, conversational intent routing, and deal-health and lead-scoring pipelines — with per-task model routing so each step runs on the model best suited to it.

Salesforce · HubSpot · Zoho
Knowledge base · multi-tenant

Every answer traces back to a source

A three-stage agent compiles each ingested document into cited, cross-linked pages, and an index-guided retriever answers questions from only the pages it needs — so nothing is asserted without a citation.

3-stage compile → index-guided retrieval → cited answers
ContextGraph · knowledge graph

One graph of the customer across four systems

Ingests Slack, Salesforce, Zendesk and Shopify, extracts entities through a combined rule-based and LLM pipeline, resolves identities across systems, and opens the result to conversational query.

Slack · Salesforce · Zendesk · Shopify
Speech · evaluation

Chose the production transcription model on evidence

Led the Arabic and English speech-recognition evaluation: three candidates benchmarked on word and character error rate across standard corpora and real customer dictations. The incumbent turned out to vary run to run and intermittently drop clauses.

11.5–15.7%incumbent WER,
run to run
Coaching agents · DSPy MIPRO

Nine agents, optimised prompts, measured lifts

Replaced hand-written prompts with automated prompt optimisation across nine coaching agents, scored against the originals.

+35%personalisation
+28%actionability
+31%conflict resolution
Voice · documents

A voice interviewer and a legal-translation agent

The voice interviewer for HumAI’s recruitment platform, and the legal-translation agent for Mairit — contracts in PDF, Word and scanned form taken through recognition into structured, reviewable output.

Realtime speech · OCR · structured review
02

Publication

Submitted to ICEAI 2026

Instrumentation-Free Continuous Monitoring of Deployed LLM Applications Using Auto-Generated Grounded Benchmarks

Z. Khan, M. Shah, S. Ali, M. S. Khan

OBAM-AI treats a deployed language-model application as an opaque HTTP endpoint: it builds a benchmark grounded in the project’s own documents, replays it against the live service on a schedule, and scores every reply for groundedness and helpfulness — with no SDK, instrumentation or code change inside the system under test. The working implementation was validated over a 24-day scheduled run against a live deployment.

Began as my final-year project at UET Mardan, where it was built and run end to end.

24d
of continuous, scheduled evaluation against a live deployment
0
lines changed inside the system under test
methodproject corpusgrounded benchmarkscheduled replayLLM-as-judge scoring
03

Timeline

2021 – 2026
Jan 2025 – present Dubai, UAE

AI Software Engineer — HumAI

Huscribe’s CRM reasoning layer, the multi-tenant knowledge base, ContextGraph, the Arabic and English ASR evaluation, the DSPy-optimised coaching agents, the voice interviewer and the Mairit legal-translation agent — detailed above.

Sep – Dec 2024 UAE, remote

AI Developer — MakTek AI

Fine-tuned CodeStral-22B to generate error-free PineScript trading indicators, paired with a sandboxed execution loop that fed compiler errors back to the model so it corrected its own output.

Built Bots For AI, a configurable chatbot platform — a knowledge base per customer with automatic content syncing, deployed across web, WhatsApp, Messenger, Instagram, Slack and Salesforce — and a WhatsApp inventory assistant answering product, pricing and stock enquiries from live business data.

Jun – Aug 2024 Islamabad, Pakistan

NLP Intern — ITSOLERA

Fine-tuned TinyBERT for sentiment classification and shipped it as a web application; built an e-commerce support assistant on FastAPI and LangChain, and an interview-preparation tool that generated questions, scored answers and returned per-answer feedback.

2021 – 2025 Mardan, Pakistan

BSc Computer Software Engineering — University of Engineering & Technology, Mardan

Final-year project: OBAM-AI, the LLM observability platform behind the ICEAI 2026 submission. Degree attested by the Higher Education Commission of Pakistan.

DeepLearning.AI certifications: Generative AI with Large Language Models · Natural Language Processing Specialization · Deep Learning Specialization · Machine Learning Specialization.

04

Stack

What I reach for
Languages
Python, TypeScript, SQL, C/C++, Dart (Flutter)
LLM & agents
LangGraph, LangChain, DSPy (MIPRO), CrewAI, LiveKit, Pipecat, MCP, multi-agent orchestration, function calling, prompt optimisation
RAG & retrieval
Multimodal RAG, hybrid search, re-ranking, ChromaDB, FalkorDB, knowledge graphs, entity resolution, index-first navigation
Training & evaluation
PyTorch, TensorFlow, Hugging Face, TRL / PEFT / LoRA, RAGAs, LLM-as-judge, WER and CER scoring (jiwer), regression suites
Backend & data
FastAPI, Pydantic v2, MongoDB, PostgreSQL, Redis, WebSockets, Socket.IO, Pub/Sub, microservices
Cloud & platform
GCP (GKE, Cloud Run, Cloud Build, Cloud Deploy, Vertex AI), AWS, Docker, Kubernetes, GitHub Actions, Pytest, Ruff, Bandit
Speech & multimodal
Deepgram, ElevenLabs, Whisper, TTS and STT pipelines, voice agents, vision-language extraction
06

Contact

Email is fastest

Let’s talk about agentic systems, evals — or a role.