Skip to content

Kathmandu, Nepal — --:-- NPT

Open to AI engineering roles

Er. PrashantPhuyal

I build the parts of a product that have to understand something — chatbots that answer from a company's own documents, and the generative AI behind them.

Prashant Phuyal, AI Engineer
1+
Year at Dome Infosys
9
AI domains shipped
8
Client AI assistants
3
Languages supported
RAGFastAPIQdrantLangChainGeminiOpenAILangGraphFAISSPostgreSQLRedisMongoDBCeleryPyTorchTensorFlowFaster R-CNNYOLOOpenCVEasyOCRLiveKitPydanticRAGFastAPIQdrantLangChainGeminiOpenAILangGraphFAISSPostgreSQLRedisMongoDBCeleryPyTorchTensorFlowFaster R-CNNYOLOOpenCVEasyOCRLiveKitPydantic

Selected work

04 projects — 2025/26

These are team products. Each one names the part I engineered and who built the rest.

The whole design is about refusing to answer rather than answering well.

Fertility treatment is a subject where a chatbot that improvises is worse than no chatbot at all. A question passes an intent classifier, a domain check and a safety check before anything is retrieved, so an out-of-scope question is turned away before it ever reaches the model. Putting that in the prompt would have been a request; putting it in front is a gate.

Answers come from clinic-approved material retrieved out of Qdrant and fed to Gemini with a prompt constrained to IVF. The knowledge base is built from whatever the clinic already had — PDFs, pages on their website, a YouTube channel, spreadsheets — so there's a parser for each rather than a requirement that someone retype it all.

It also takes phone calls. The voice agent runs on LiveKit, and the part that took longest was barge-in: letting the caller interrupt mid-sentence and having the agent stop talking and listen. Without it you have an IVR menu; with it you have something closer to a conversation.

Role

I built the AI layer — retrieval, the voice agent and the service behind them. The mobile app and surrounding platform were built by fellow developers.

Built with

FastAPIQdrantGeminiOpenAILiveKitPostgreSQLRedis

Nine features, one router. None of them know which model answered.

This one is a separate service rather than a feature. Different teams owned the CRM, the booking system and the customer app, and all three needed AI — so instead of building it into any one of them, I built it as its own service the others call over HTTP.

It covers nine areas: the chatbot, voice calls, damage detection from photos, price estimation, provider matching, provider scoring, review sentiment, fraud and fake-review detection, and analytics write-ups. They vary a lot, but anything involving a model goes through one router — none of them import an LLM SDK directly. Gemini is primary and OpenAI picks up automatically when it fails, and because that switch lives in one place, nine features didn't have to care.

The chatbot works in Nepali first, then English and Hindi. The awkward part is that people write Nepali in Latin letters with no agreed spelling, so a word list for detecting it is hopeless — the same word turns up five ways. It looks at word structure instead.

Role

I built the AI service. The CRM, booking and customer-facing products were built by fellow developers.

Built with

FastAPIQdrantFAISSGeminiOpenAILangChainPostgreSQLRedisCelery

It's published under a real person's name, so the AI doesn't get the last word.

The platform turns interviews, documents and photographs into a book or a biography. The constraint that shapes everything is that it's published under a real person's name, so the AI isn't allowed the final say. It extracts and it drafts; a human checks what's true in between.

Sources are transcribed or read with OCR, then pulled apart into facts. Two people remembering the same event differently is normal in a biography, so the knowledge module looks for contradictions and surfaces them rather than quietly picking one. Only after someone has verified the material does anything get drafted into chapters.

It's also the first thing I've built with a proper evaluation setup — test cases in English and Nepali with metrics attached, so when I change a prompt I can tell whether it got better or just different. I should have been doing that earlier.

Role

I built the AI engine. The platform around it was built by fellow developers.

Built with

FastAPICeleryQdrantGeminiPostgreSQLRedisPytest

A score alone is useless. The explanation is the product.

Two sides of the same problem: employers write thin job posts because writing a full one is tedious, and candidates get filtered out by applicant tracking systems without ever being told why.

For employers, a few words about a role expand into skills, keywords, responsibilities and a full specification. For candidates, the CV maker writes a summary, suggests skills, scores the CV the way an ATS would and explains what's dragging the score down — then generates interview questions from what's actually in their CV.

None of it is a chat window, so the output has to be valid structured JSON every time rather than prose that happens to parse. That turned out to be most of the work.

Role

I built the AI APIs. The web platform and mobile apps were built by fellow developers.

Built with

PythonFastAPILLM APIsStructured JSON

Also built

Personal & academic

Try it

Ask this site a question.

A retrieval engine running in your browser over this page's own content. It shows you the passages it found and what it scored them — and when nothing scores highly enough, it says so instead of guessing.

About

Kathmandu, Nepal

I'm an AI engineer at Dome Infosys in Kathmandu, where I started as an intern a little over a year ago and now work across our clients' AI features.

Most of my work is retrieval-augmented generation, which in practice is less about prompting than about everything around it: parsing the files people actually have, chunking them so a search can find the right passage, and deciding when the retrieved context is too thin to answer at all. The last one matters most. A model that invents a plausible answer about IVF treatment is worse than one that says it doesn't know.

Before that I spent four years on a computer engineering degree, most of it on computer vision — object detection and OCR for Nepali, which has far less training data available than English does.

Role
Associate AI Engineer
Company
Dome Infosys
Since
More than 1 year · Promoted from AI/ML Intern
Degree
Bachelor in Computer Engineering
University
Cosmos College of Management & Technology, 2021 – 2025
Result
CGPA 3.6 / 4.0
Certified
Data Science with Python
Languages
Nepali (Native), English (Professional), Hindi (Conversational)

Experience

More than 1 year

Dome Infosys

Associate AI Engineer

Promoted from AI/ML Intern

Clients served

Vatsalya IVF (Ziva)Babyloan SchoolYatriFlyNepwoodMakaluSewaFundRojina Beauty SalonPokhara Motors

I design, build and deploy production RAG chatbots and generative AI features for clients across healthcare, education, travel, recruitment, retail and financial services.

  • 01Own end-to-end RAG pipelines — ingestion, chunking, embeddings, Qdrant and FAISS indexing, top-K retrieval, and grounded generation with domain-specific prompts.
  • 02Build asynchronous FastAPI backends with conversation persistence in PostgreSQL and MongoDB, Redis caching and rate limiting, and automated deployment through GitHub Actions.
  • 03Delivered document-grounded assistants for eight clients, from a fertility clinic to a school to a motors dealership.

Capabilities

As grouped on my CV

Generative AI & LLM

RAGPrompt EngineeringAI ChatbotsVoice AI AgentsIntent ClassificationConversation MemoryStructured JSON OutputLangChainLangGraphGeminiOpenAIHuggingFace

Vector Search & Retrieval

QdrantFAISSEmbeddingsSemantic SearchDocument IngestionChunkingTop-K Retrievalsentence-transformers

Backend & APIs

PythonFastAPIAsync PythonPydanticREST DesignSQLAlchemyCeleryJWT AuthRate LimitingGitHub ActionsPytest

ML & Computer Vision

PyTorchTensorFlowKerasScikit-learnFaster R-CNNYOLOVision LLMsOpenCVEasyOCR

Databases & Data Analysis

PostgreSQLMongoDBRedisSQLPandasNumPyMatplotlibEDA

Contact

Let's talk.

Open to AI engineering roles, and happy to talk about retrieval, grounding, or getting a model to work in a language it wasn't really built for.