Rajarshi Haldar Contact

PhD Candidate · Computer Science

Rajarshi Haldar

University of Illinois Urbana-Champaign

I study whether language models can be trusted to judge other language models.

  • LLM Evaluation
  • Interpretability
  • Code Intelligence
  • NLP
Rajarshi Haldar

About

Hi, I am a PhD candidate in Computer Science at the University of Illinois Urbana-Champaign, advised by Prof. Julia Hockenmaier (opens in a new tab) in the Hockenmaier Lab (opens in a new tab). My work asks what a language model represents while it grades another model, and whether the score it gives can be trusted.

LLM judges can assign different ratings to identical responses, making a single benchmark score an unstable measurement. I first studied that problem from the outside, then turned to the models' internal activations with sparse autoencoders. I found that it is possible to track human expert annotations with only a few dozen features, as simple probes fitted to those features predict human ratings better than the judge's own output does. Therefore, the LLM judges contain more information about the quality of an item than their outputs can communicate.

I also study how language models represent source code. In my work on code search and code summarization, I ask which features models use when they try to understand the logic of a function. During research internships at IBM Research I worked on event extraction and schema induction, and that work led to three granted US patents.

Before Illinois, I got my B.Tech. in Computer Science and Engineering from the Indian Institute of Technology Kharagpur.

Research Interests

The questions below run from the score a judge gives to the structure a model recovers from text.

  • Can an LLM judge be trusted?

    I measure self-inconsistency in LLM-as-a-judge frameworks and ask what unreliable scores mean for benchmarks built on top of them.

  • What does an evaluator represent internally?

    I use sparse autoencoders and probing to isolate features that track generation quality, then test whether a rating is faithful to what the model represents in its residual activations.

  • What does a model notice in source code?

    I study semantic code search, code summarization, and the features models capture while they match natural language to the logic of a function.

  • How can structure be recovered from text?

    I use event schema induction, event extraction, and knowledge graphs to recover structure from unstructured text so it can be searched, compared, and reasoned over.

Publications

Selected work on LLM evaluation, code intelligence, and structured knowledge. The full and continuously updated lists are on Google Scholar (opens in a new tab), DBLP (opens in a new tab), and the ACL Anthology (opens in a new tab).

  1. Consistently Good vs. Occasionally Great: A Rubric for Open-Ended Feedback Quality from Humans and Machines

    Preprint, arXiv:2608.21850

  2. Rating Roulette: Self-Inconsistency in LLM-As-A-Judge Frameworks

    Findings of the Association for Computational Linguistics: EMNLP 2025

    The same judge scores identical inputs differently from one run to the next, with Krippendorff's alpha as low as 0.27.

  3. Analyzing the Performance of Large Language Models on Code Summarization

    Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING)

    Code summarization scores can follow token overlap in function names more closely than the structure of the code itself.

  4. A Multi-Perspective Architecture for Semantic Code Search

    58th Annual Meeting of the Association for Computational Linguistics (ACL)

    A multi-perspective architecture lifted MRR on semantic code search by 28% over a code-token baseline.

  5. Deep Learning Driven Venue Recommender for Event-Based Social Networks

    IEEE Transactions on Knowledge and Data Engineering, 32(11)

  6. A Validated Scoring Rubric for Explain-in-Plain-English Questions

    51st ACM Technical Symposium on Computer Science Education (SIGCSE)

  7. Cluster Aware Mobility Encounter Dataset Enlargement

    15th International Wireless Communications & Mobile Computing Conference (IWCMC)

  8. CL Scholar: The ACL Anthology Knowledge Graph Miner

    2018 Conference of the North American Chapter of the ACL: Demonstrations (NAACL-HLT)

Patents

Granted US patents from internships at IBM Research, on event extraction, schema induction, and code search. All are assigned to International Business Machines Corporation.

  1. Unsupervised Event Extraction

  2. Graph-Based Event Schema Induction for Information Retrieval

  3. Multi-Perspective, Multi-Task Neural Network Model for Matching Text to Program Code

Experience and Education

Research across language models, code understanding, and applied NLP.

  1. PhD Candidate and Graduate Research Assistant

    University of Illinois Urbana-Champaign · Urbana, IL

    Advised by Prof. Julia Hockenmaier

    • Probes beat the judges' own scores on 7 of 8 criteria when fitted to sparse autoencoder features (in review).
    • Showed LLM judges rescore identical inputs inconsistently, with Krippendorff's alpha as low as 0.27 (EMNLP 2025).
    • Traced LLM code-summarization scores to token overlap in function names, not code structure (LREC-COLING 2024).
    • Introduced a multi-perspective code-search architecture, lifting MRR 28% over a code-token baseline (ACL 2020).
    • Built the dataset, training, and evaluation pipelines in PyTorch and Hugging Face Transformers.
  2. Digital Intern

    Schlumberger · Menlo Park, CA

    Software Technology Innovation Center (STIC)

    • Built a knowledge-graph embedding model that learns from limited labeled data via labeling functions and Snorkel AI.
    • Shipped a query-answering model on those embeddings that replaced the team's system at lower latency and overhead.
  3. Graduate Research Intern

    IBM Research · Yorktown Heights, NY

    • Led an automated tech-support agent that resolved IT queries by extracting events from unstructured incident text.
    • Raised retrieval MRR from 0.468 to 0.538 by inducing schemas over 802K technical documents.
  4. B.Tech. in Computer Science and Engineering

    Indian Institute of Technology Kharagpur · Kharagpur, India

    • Undergraduate research on knowledge graphs over the ACL Anthology (NAACL-HLT 2018 demo) and on venue recommendation for event-based social networks (IEEE TKDE).

Teaching

  1. Teaching Assistant, CS 447 — Natural Language Processing

    University of Illinois Urbana-Champaign · Urbana, IL

    • Designed a Google Colab coding assignment on machine translation and wrote its auto-grader on Gradescope.
    • Held office hours twice a week and answered student questions on Campuswire.
  2. Teaching Assistant, CS 416 — Data Visualization

    University of Illinois Urbana-Champaign · Urbana, IL

    • Graded a midterm and a final project in which students built dashboards in Tableau and D3.js.
    • Held weekly office hours and answered student questions on Campuswire.

Honors and Service

Get in touch

Happy to chat about LLM evaluation, interpretability, code intelligence, or anything adjacent.

Privacy

How this site handles your data

This is a personal homepage. It does not collect anything about you, nor does it build a profile of anyone who reads it.

Who is responsible

I, Rajarshi Haldar, run this site as an individual, not on behalf of any employer or institution. Where applicable, I process personal data in accordance with the GDPR and the UK GDPR. I am the controller for personal data that I keep.

The Email me button on the homepage is the quickest way to reach me. If you do not have a mail client, ask through any other route you already have (a paper's corresponding address or a department directory).

What is collected

Nothing is collected. There is no script that measures your visit, and no report about it is sent anywhere.

There is no advertising, no data broker, and no third-party tag. Nothing here follows you to another site.

What is stored on your device

Nothing is stored on your device. There are no cookies set, and nothing is written to local storage.

The default color theme will be determined by your system setting on every visit. Since I do not store your preference for this website, switching it with the toggle lasts only until you reload, after which it reverts to the default theme.

What gets loaded on the pages of this website

Every file comes from this domain. The typefaces are served from here rather than from a font network, so reading a page does not contact any other company.

Who else is involved

Cloudflare, Inc. hosts the site. Every request reaches its servers before it reaches these pages, so the request data any host keeps (an IP address, a time, the page asked for, the browser that asked) is created there. I have never seen it and I do not hold any copy; none of it is ever shown to me.

Cloudflare provides hosting, delivery, and security services. Depending on the service and data involved, Cloudflare may process request data as a processor for me or as an independent controller. I do not access visitor request logs. Cloudflare describes its roles and retention practices in its privacy policy (opens in a new tab). The applicable service and data-processing terms identify the relevant role and retention period.

Cloudflare and Google may process personal data outside the EEA or the United Kingdom. Where a restricted transfer occurs, I rely on an applicable adequacy decision, including the EUUS Data Privacy Framework or UK Extension, only when the recipient has active coverage.

Email basis

The only data I keep is correspondence I receive, and the basis for keeping it is legitimate interests under Article 6(1)(f), specifically the interest in answering professional and academic correspondence.

There is no automated collection and no fixed retention schedule. I keep correspondence as long as it serves the purpose of the exchange, such as answering a question, managing a matter, or keeping a record of what was discussed. On request, I erase specific correspondence where the GDPR conditions for erasure apply, except where retention is necessary to comply with a legal obligation or to establish, exercise, or defend legal claims.

Email

The Email me button opens a message addressed to a Gmail account. If you send the message, Google receives it as the mailbox provider, and its own privacy policy (opens in a new tab) governs what it does in that capacity.

What I hold is your address together with the contents of the message you send.

Your rights

You may request access to, correction of, erasure of, or restriction of processing of your personal data where the GDPR conditions apply. You may object at any time to processing based on legitimate interests. Data portability applies only where the conditions in Article 20 are met. Those conditions normally do not apply to correspondence processed under legitimate interests. These rights have legal exceptions.

Ask through the Email me button or by any other route you have for reaching me.

You may also complain to your national supervisory authority. In the United Kingdom, that is the Information Commissioner's Office (opens in a new tab). In other countries, it is the authority designated by that country's data-protection law.

Automated decisions

None. Nothing here makes a decision about you by automated means or builds a profile about you.

Changes

Any revision updates the date on this page.

Last updated September 7, 2026