
- Early StageStartup in initial stages
AI Intern
- ₹10,000 – ₹12,000 • No equity
- |Remote ()
- |No experience required
- |Internship
About the job
About the role
Legal work is language work. Contracts, trademark filings, judgments, and templates are dense, long, adversarially drafted, and unforgiving of small errors — which makes them one of the hardest and most interesting domains to apply modern NLP to.
We're looking for a third-year engineering student to join us as an Applied AI Intern. This is a balanced research-and-build role: roughly half your time reading, prototyping, and measuring whether a technique actually works, and half turning what works into something our customers can use. You will not be handed a fully-specified ticket queue. You'll be given a problem, a dataset, a quality bar, and a lot of room.
The field moves faster than any curriculum can keep up with. We care much more that you can read a paper or a model card, form your own opinion, and test it against real data than that you already know any particular framework.
Problems you could work on
- Structured extraction from messy documents — clause- and entity-level extraction from contracts, filings, and scanned records, where layout, tables, and OCR noise all fight back.
- Document comparison and redlining — semantically meaningful diffs between document versions, not just character-level ones.
- Template and draft generation — generating first-draft legal documents that are grounded, consistent, and safe to hand to a lawyer.
- Retrieval over large legal corpora — chunking, hybrid retrieval, reranking, and citation grounding on documents far longer than any context window is comfortable with.
- Making LLM features reliable — reducing hallucination, enforcing structured outputs, handling failure modes, and keeping latency and cost inside a budget.
- Evaluation infrastructure — the unglamorous work that makes all of the above measurable rather than vibes-based.
What you'll do
Research and prototyping
- Track developments across NLP, LLMs, retrieval, and agentic systems; separate genuine advances from hype, and say which is which.
- Read papers, model cards, and technical write-ups closely enough to reproduce the core claim on a small scale.
- Build fast, throwaway prototypes to answer a specific question, and be willing to kill them when the answer is "no".
Building and shipping
- Take a validated prototype to a working feature: clean interfaces, sane error handling, reproducible pipelines.
- Build retrieval and extraction pipelines end to end — parsing, chunking, embedding, retrieval, reranking, grounding.
- Design prompts and context strategies as engineering artifacts: versioned, tested, and reviewed, not pasted into a notebook.
- Where a smaller task-specific model beats a general one on cost, latency, or accuracy, fine-tune or adapt one and prove the trade-off with numbers.
Measurement and evaluation
- Build and maintain evaluation sets — including adversarial and edge cases drawn from real legal documents.
- Choose metrics that reflect what actually matters to the user, and be honest about what they miss. Automated judging is useful and is not a substitute for human review; calibrate one against the other.
- Track regressions, quality, latency, and cost together. A change that improves accuracy by 1% at 10x the cost is not obviously a win.
- Run experiments with enough statistical care that the result survives scrutiny.
Data
- Collect, clean, label, and version datasets — often the highest-leverage work on the list.
- Handle client data with appropriate care around confidentiality, privilege, and PII. This is non-negotiable in legal tech.
Communication
- Document what you built, what you tried, what failed, and why. Negative results written up clearly are genuinely valuable to us.
- Present findings to engineers, product, and domain experts, adjusting the depth for the audience.
- Work with lawyers and domain experts to understand what "correct" means before optimizing for it.
What we're looking for
Core
- Currently in the third year of a B.E./B.Tech/integrated M.Tech in Computer Science Engineering, or a related field.
- Strong Python. You can write code others can read, debug, and build on.
- Solid NLP fundamentals: tokenization, embeddings, similarity and retrieval, sequence labelling, and the evaluation metrics that go with them.
- A working understanding of transformer-based language models — how they're trained, why they fail, what context length and attention actually cost you.
- Practical experience with at least one deep learning framework, and enough comfort with the standard Python data and analysis stack to explore a dataset without hand-holding.
- Experiment design and statistical literacy: sample sizes, baselines, confounders, and why a single benchmark number rarely settles an argument.
- Version control fluency (Git) and the habits that come with collaborative development.
- Clear written and verbal reasoning. You can defend a position and also change your mind when the evidence moves.
Strong signals (not requirements)
- Projects you've built with LLMs — RAG systems, agents, extraction pipelines, evaluation harnesses. Shipped and imperfect beats polished and theoretical.
- Experience fine-tuning or parameter-efficient tuning of open-weight models, and a sense of when it's worth it.
- Familiarity with vector or hybrid search, rerankers, and the failure modes of each.
- Working with long, structured, or scanned documents — PDF parsing, layout-aware extraction, OCR cleanup.
- Building evaluation harnesses or benchmarks of your own.
- Open-source contributions, technical writing, or peer-reviewed publications.
- Data visualization, and enough web/API development to put a prototype in front of someone.
- Any exposure to the legal, compliance, or regulatory domain.
How you work
- Self-starting and comfortable with ambiguity — the problem statements above are deliberately underspecified.
- Detail-oriented. In legal tech, a plausible-sounding wrong answer is worse than no answer.
- Collaborative, communicative, and good at managing your own time.
- Intellectually honest about what your model does and doesn't do.
How we'll evaluate you
Projects are the fastest way for us to understand what you can do, so send them. A GitHub repo, a write-up, a demo, or a short note on something you tried that didn't work — all count.
What we look for in a project: a clearly stated problem, an honest account of the approach, some evidence you measured the result, and a sense of what you'd do differently. We're more interested in your reasoning than your leaderboard position.
The process itself is two stages. First, we send you a task — a scoped problem close to the kind of work you'd actually do here. Then a technical round, where we go deep on what you built: the choices you made, the ones you rejected, and how you'd extend it.
About the company

MikeLegal
- Early StageStartup in initial stages
Employees joined from
Similar Jobs





