Student-led Network: LLM Learning Group

The focus of the network is dedicated to fostering an inclusive, collaborative community for individuals interested in the practical and methodological integration of large language models (LLMs) within social sciences research. As LLMs rapidly emerge as a novel and powerful tool for generating, analysing, and interpreting social data, this network provides a space to critically and creatively engage with this avenue. Our primary focus is to create a supportive environment that encourages knowledge sharing, peer learning, and collective exploration of innovative uses of LLMs in both research and daily activities.

Through these activities, the network seeks to build capacity, lower barriers to engagement with LLM technology in research applications, and strengthen the ongoing development of best practices in social science research.

If you are interested, you can:

Group picture of PGR's attending a LLM Learning Group Event

Events Recap:

Establishing the Network:

The first LLM Learning Group meeting was held online on 10th April and was planned for one hour. Around 10 PhD students attended, representing a broad range of disciplinary backgrounds, including social sciences and computer science.
The session was jointly facilitated by Chin Wei, Jacob, Lu Ping and Yao, primarily to establish the shared goals/purposes of the LLM Learning group and to give participants an opportunity to meet each other virtually. Participants introduced their doctoral projects and discussed their prior experiences using Large Language Models (LLMs) in both research and everyday work contexts.
Discussion focused on how LLMs might support social science research, alongside limitations and risks involved in their use. Some Topics raised included the reliability and possible fabrication of LLM-generated information; risks to data confidentiality when working with sensitive research material; bias in model outputs; disclosure of LLM use in academic writing; and the difference between using an LLM as a practical research aid and treating it as a valid research method.
Jacob then subsequently delivered a short presentation of about 20 minutes on his research area, which concerns the use of LLMs with public health data. This gave participants a concrete cross-disciplinary example of how LLMs is being applied to complex real-world data.

Prompt Engineering Workshop:

The second LLM Learning Group meeting took place online on 27th May and was also planned for one hour. Seven PhD students attended, bringing a range of disciplinary backgrounds and levels of previous experience with LLMs for interdisciplinary discussion.
The organising team prepared and delivered a session on prompt engineering, focusing on how the wording, structure and context provided in a prompt can influence an LLM’s response. A short presentation introduced the structure of the prompt, then demonstrated techniques such as setting explicitly goals and constraints, assigning the model a role, using zero-shot vs few-shot examples, prompting step-by-step reasoning through chain-of-thought, and self-consistency through multiple reasoning paths.
Next, participants then took part in an interactive activity designed to demonstrate the effect of prompt design on output quality. By using prompts to analyse a qualitative dataset of news articles related to Hyde Park and Headingley in Leeds, including judging local relevance, identifying topics and problem framing, and examining whose voices were represented.
Subsequently, the discussion covered possible uses of prompt engineering in PhD work, including generating structured outlines, refining research communications, summarising non-sensitive material, supporting early-stage idea development or creating reproducible instructions for LLM-assisted tasks. Participants also considered the need to review outputs critically, particularly where accuracy, disciplinary context or ethical judgement matters.

Learnaway Day:

On 20th July 9:30am t0 3:30pm we hosted an in-person Learnaway Day, held at the University of Leeds. Fifteen doctoral researchers attended from institutions across the WRDTP, from a whole range of academic disciplines. This larger face-to-face event created time for sustained discussion, informal networking and peer exchange across institutional boundaries.
The programme centred on student presentations in the morning about their use of LLMs in social science research. By giving students space to present their own interests, ideas and emerging work, the event supported the network’s aim of sharing knowledge through peer-led contributions rather than relying only on organiser-led teaching. It also exposed attendees to different disciplinary perspectives on how LLMs may be used in research.
The organising team also delivered an interactive activity comparing synthetic data generated using LLMs with real data. Participants considered similarities and differences between the two, including the extent to which synthetic data may preserve meaningful patterns, the risks of reproducing or amplifying biases, and the limits of treating generated data as a substitute for data collected from real people or settings.
Additionally, this activity prompted a thoughtful discussion about the purpose and value of synthetic data in social science research. Participants questioned when synthetic data might be appropriate, such as when it could create misleading conclusions because it lacks the complexity, uncertainty or social context of real-world data.