The EMMA Workshop, a JT-DH pre-conference event, will take place on September, 16th, from 9.30am to 1pm, at the Faculty of Computer Science.
The event is free to attend.
Contact: Jaya Caporusso, jaya.caporusso@ijs.si
Natural Language Processing for News Analysis: Theory and Practice – Results and Experiences from the EMMA project
ABSTRACT
This workshop presents the results and experiences of the EMMA project and provides a comprehensive overview of modern natural language processing methods for automated news analysis. In the first part, participants will be introduced to the EMMA project, which brings together the Jožef Stefan Institute, the Faculty of Computer and Information Science at the University of Ljubljana, and the industry partner Kliping. The workshop will address the real-world needs and challenges of media monitoring. A related project focusing on advanced approaches to information retrieval and analysis will also be presented. The second part focuses on specific methods for news analysis, including keyword extraction, named entity recognition, topic classification according to the IPTC standard, sentiment analysis, grouping articles into stories, automatic summarisation, and diachronic analysis of media discourse. Participants will learn about key research approaches, results, and practical examples of multilingual language technologies, as well as their potential applications for monitoring large volumes of media content and analysing changes over time in the meanings, topics, and narratives found in media coverage.
PROGRAMME
Introduction
9h30-10
Senja Pollak: Presentation of EMMA project
Sebastjan Hribar: Media monitoring: real life needs and challenges
Emma results
10-11 Chair session Senja Pollak
- Nikola Ivačič: Client-specific news classification and multilingual NER
- Marko Pranjić: Adding Context to Keyword Extraction
- Nikola Ljubešić: Building a multilingual IPTC topic classifier on a budget – LLM data annotation and cross-lingual models to the rescue
- David Blazheski: IPTC topics – multi-label fine-grained classification
11h-11h30: Coffee break
11h30-13h Chair session Nikola Ljubešić
- Matthew Purver: Document-level sentiment analysis in news articles
- Nishan Chatterjee: Document-level aspect-based sentiment analysis in South Slavic News
- Tadej Škvorc: Discovering stories across news articles
- Marko Robnik Šikonja: Quality, Length-Compliance, and Practical Deployability in News Summarization
- Marko Pranjić: Diachronic Semantic Analysis of Text Corpora
