
On 2 September, the SOAD community met for their first hackathon on open academic data. The theme was inspired by another notable event in Switzerland, happening just the day after: The award ceremony for the IG Nobel Prize.
True to their motto “Research that makes you first laugh, and then think”, the 20 participants picked their favourite topic from a selection of IG Nobel Prizes to investigate further on. Among the topics: Why does coffee spill when you walk while carrying a cup? Do chicken walk like dinosaurs when they are fitted with artificial tails? And can a dead spider be turned into a working mechanical gripper? The task: Find out as much as possible about the topic using bibliometric tools. This was purposely kept broad to give free rein to creativity. The tools used during the hackathon also reflected the different coding skill levels of the participants, ranging from no experience to expert coder.
Most users knew OpenAlex already and this was also the most frequently used resource in the hackathon. That included the web-based semantic search as well as a more code-heavy approach, using the API and an LLM-supported python script to quickly scan the retrieved data. Thankfully, OpenAlex supported us by increasing the daily limit of queries. This allowed the participants to dig deep into exploration without running out of credits.
The comprehensiveness of OpenAlex allowed broad overviews on the topics chosen, such as publications with similar research questions, following or preceding the IG Nobel awarded paper. While title and keyword-based searches, using the semantic search, delivered good results, several groups observed that items tagged as related were not really related or too far away from the original thought. There, citation-based tools such as connected papers and research rabbit delivered more relevant results. Both do not require any coding skills from the users to already create insights. Some groups used python code to derive and analyse data through the OpenAlex API rather than browser-based results only. Specifically, this group explored an IG Nobel Prize which was awarded for “for determining by experiment whether it is safer to transport an airborne rhinoceros upside-down”. From their broader exploration around the topic, they found the surprising insight that cats appear to be most frequently airlifted mammals – at least based on the number of publications in OpenAlex containing the word “airlifted” and a species’ name.

Picture 1: Number of articles on airlifted species by countries, for the species with most articles
Still, working with (open) academic data can prove “messy”, as one participant stated. Bibliometric inconsistencies leave a certain error level in the data. While misattributed language might be a minor problem for many use cases, misattribution of affiliation may blur results for institutional reporting. Not surprisingly, one of the major wishes of the participants towards open data is an improved level of data quality – acknowledging at the same time that this does require community effort. And we are happy that the Swiss Open Academic Data community can be one of the platforms to foster this.

Picture 2: Wishes for open data from the participants: “Improved data quality, of course (it’s up to us!)” and “Better data quality through community curation”
Besides meeting new people and discovering new tools, one participant voiced what might be the most important take home of the day: “Open data let’s you tell great stories!”
Keep it open and stay tuned for the upcoming spring event!
