Citation is Collaboration: Software Recognition in Research and Industry
In many fields, including research and industry, software is essential for driving innovation and scientific discovery, yet mechanisms for crediting software developers remain inconsistent and underdeveloped. This lack of recognition, particularly for open-source contributions, can discourage participation in software development and limit career opportunities for developers. Astronomy, as a computationally intensive discipline with a rich history of open-source software contributions, offers a valuable case study to examine these challenges. Over the past decade, changes in journal policies and emerging publishing practices have sought to address the issue, but their impact on credit attribution remains unclear.
This talk addresses the issue of software credit by analyzing publication and citation practices in astronomy. It evaluates how existing policies acknowledge software contributions and examines variations across journals and over time. Drawing on bibliometric data from the past decade, the analysis focuses on citation patterns for commonly used libraries, trends in citation rates, and the influence of journal policies. The study includes both foundational libraries, such as NumPy, and astronomy-specific libraries, such as Astropy. Based on these findings, the talk will offer recommendations to enhance the attribution of software contributions.
The issue of recognizing research software is not unique to astronomy or research. Participants from industry, other computationally driven fields, open-source communities, and publishing will find the insights applicable to their own disciplines. Understanding how software is cited and credited is critical for shaping more equitable recognition systems, which in turn support sustainable software development and community growth. The audience will leave with a clear understanding of how astronomy’s experience can inform broader efforts to address similar challenges in their respective fields.
The data and code for the analysis will be shared with participants giving participants access to a reproducible framework for analyzing software citation practices in other disciplines or software ecosystems.
This session took place in track Research Software Engineering and was classified suitable for novice domain / novice python by the speaker.
Transcript (auto)
Auto-generated from the recording utilizing Open-Source AI. Speaker labels (Speaker 1, Speaker 2) reflect diarization, not identity. Timestamps refer to the recording.
Speaker 1 [00:08]
Hello everyone, thank you for joining us for this talk. So let me just briefly introduce myself at the start. My name is Iva Momtcheva and I work at the Max Planck Institute for Astronomy in Heidelberg. And for this talk specifically, I wear three hats. First, I'm a research astronomer. I do research on galaxy evolution and I publish scientific research in this area. Also, I'm the head of the data science group at the Max Planck Institute for Astronomy, where I work with researchers and I help them develop open-source software and I help them also publish this software. And the third hat that's relevant for this talk is I'm an editor for the Journal of Open Source Software, where I help authors from around the world publish papers on their software so they get credit for their work. So with that in mind, let me start with a brief motivation. So I'm sure you are all familiar with this XKCD comic showing kind of schematically how modern day software infrastructure looks like. I'm sure you have worked on or your software has depended on a project that looked like this. There's always, there's a stack of dependencies all the way down and there's this long list of people who maintain and develop these dependencies. Like this guy in Nebraska who has been maintaining this small project that this whole pile depends on. But for the purposes of this talk, I wanna ask what motivates this person and how can he get credit for this work? And not just him, but everybody else in the stack, maybe this project is trying to get a grant that supports their work. How can they show that their project is important? Maybe this project is supported by a junior researcher who needs to get credit so that he can get his next job. And maybe this project here is supported by Numfocus and they want to show that the work that they're funding is useful. And not just this, but how can we recognize all of the efforts in the stack? So to summarize, we all care about three things. Jobs. We want to get our next job. So how can we get credit for the work that we have done and what kind of credit matters for not just for yourself but also for the other contributors to the project. What kind of credit is recognized for your current employer and also for your future employer. We all care about money for our projects. What kind of credit is important for your current funders so that they appreciate that your project is useful and also what kind of credit is useful for your future funders so that you can get your next grant or your next support. And finally, we care about sustainability. What information is needed to reproduce results from software, but also how can you ensure that this whole stack of projects that are underlying your infrastructure is alive and well in the long term so that your own project can continue being useful for your users. So software is critical for the advancement of much, lots of things, for everything at this conference, but also for science specifically, in the context of the rapidly growing datasets that we're experiencing, not just in astronomy, but also in many other areas. However, it's widely acknowledged that software has not been consistently cited, despite its importance. And it's actually very difficult to find information on software citation. One of the pieces of information I found was the study from 2016, where the author specifically looked in great detail at only 286 papers and counted all of the citations of software in them. This was an extremely laborious study, and it was done almost 10 years ago. It's very difficult to do this kind of work consistently across many papers, but maybe Maybe we can do better and do it over again to see how citations change over time. So maybe we can do that better. So this was my motivation to do the work that I'm going to present to you today. So I'm going to use astronomy as a case study for a number of different reasons. First, I'm an astronomer, and I'm well familiar with the ecosystem. But also astronomy is a computational discipline. This is actually a plot from a paper I published 10 years ago. This is the least informative plot ever, but it shows the fraction of participants in a study who answered yes to the question, do you use software in your research? This was over a thousand participants, astronomers from around the world and across different sub-disciplines. Astronomy is also dominated by Python. This is a couple of plots from two papers, one in 2015 and another one which I'm currently working on that shows that the fraction of participants who use Python in their software development and their research has risen from 67% to 93% over the course of the last three years. So 93% of all astronomers today use Python for their research. Also astronomy has a strong open source ecosystem. It is under the umbrella of the AstroPy project, which develops the AstroPy package, but also supports a number of different libraries, that supports astronomical research, there's a really lively and large ecosystem of software projects in astronomy. I also carried the 2024 survey of software used in astronomy, which also gave us some more recent insight about the practices and what kind of libraries astronomers use. So this is from the study that shows that there are five libraries that are dominant across all astronomy researchers. Those are NumPy, Matplotlib, AstroPy, SciPy, and Pandas are used by more than 25% of all researchers who participated in this study. And finally, in astronomy results are shared via refereed papers in a common database which has a very nicely documented API which can be searched. Astronomy also has a relatively simple publishing system. Three different journals publish almost 50% of all research papers in astronomy. Those are monthly notices of the Royal Astronomical Society, published in the UK, ANA, which is published by the European Astronomical Society, Physics Reviews D, which is mainly for papers for cosmologists and particle physicists, Geophysics Review Letters, which is mainly for planetary science, and the SS journals, which is published by the American Astronomical Society. These five publications publish over 10,000 papers per year, and this is the last 13 years of publications. Back in 2011, there were just over 10,000 publications published per year across these three journals, and that has gone up to just over 16,000 last year. These five publications also map an evolving publishing landscape. They have a range of different policies and a range of different recommendations and a range of different practices around software citation. The SS journals, which are published by the American Astronomical Society, have very strongly developed recommendations and have even hired software editors to make sure that all journals, all papers submitted to them undergo software review and that they cite software properly. The geophysical review letters has had a software citation policy since 2021, but they don't actually do any enforcement on this policy. The monthly notices of the Royal Astronomical Society, which I gave two stars, just introduced a software citation policy in 2023, and the other two publications have no recommendations around software citation, and they do no enforcement around that. So with that landscape, the way I went about looking how software is cited by astronomical journals is I used the Astrophysics Data System, which is a digital library portal for researchers in astronomy and physics, and is operated by the Smithsonian. I used their API to look for mentions of the five top most commonly used libraries in the body of the papers and the bodies of the publications and the acknowledgements in all refereed astronomy papers from 2011 to 2014. The time range was determined by AstroPy, which was actually first published in 2011. All the other four libraries existed prior to that, but that just spans kind of like the same period for all five libraries I considered. I also, to kind of get a broader picture, I also used the PyPy Google BigQuery database to see kind of like what the activity is in terms of downloads for some of these packages. So first look, these are the citation counts for the five most commonly used libraries the body of the papers and in the acknowledgments of the papers over the 13-year periods that I explored. So you can see that starting 2011, there was almost no citations for these packages. And then starting somewhere in 2015, 2016, the number of citations picks up. So those are just the raw counts across all of the five journals that I looked at. Of course, the number of publications changed over time, so I wanted to normalize by that. So these are the fractions, the percentages, the fraction of papers that actually cite any of these libraries. So starting like 2015, 2016, where there was like 1%, 2%, in recent years we have peaked and kind of plateaued at about 10% to 12% across all the libraries. Of course, this is, at first glance, this is very different from a factor of five lower than the researchers that actually say that they use, self-report that they use these libraries. So there seems to be a large discrepancy between the usage of these libraries and their citation and their mention in published literature. Do the citations actually vary as a function of the policy of the journals? Doesn't seem so. So the blue line here for these plots, I'm showing the citations for NumPy and AstroPy, two different libraries, one more general and one more specific for astronomy, NumPy on the top and AstroPy on the bottom, across different journals. So the blue bars are the double-S journal, the one with the five-star citation policy, and yes, they're very good, but the geophysical review letters, which also has a very good citation policy, has almost no citations. So there seems to be a large discrepancy, and maybe that's just due to the enforcement of double-S, who does have software editors who recommend citations. But even for the double-S journals, the blue bars, the citation rate is still very low, like 15 to 20 percent actually do cite these packages, which is still much lower than the self-reported use across the field. How about the journals with the worst citation policies? Again, doesn't seem to be any dependence. The Physics Review Letters D has almost no citations in any of the papers, and the AAAS and the ANA, which is the kind of dark orange bars, has about around 10 percent, peaks at around 10 percent, which maybe also speaks that there's a difference in the field, field dependence, so one of these journals is mostly focused on publications from physics, kind of physics and cosmology, whereas the other one is more galaxy evolution and other, so even if there is no citation recommendations, there seems to be differences between different research areas. Where in the text are the citations? So the first takeaway from this figure is that, of course, most of the papers don't actually have any citations. Here I'm just showing NumPy citations in 2024 for the different journals that I explored. So the most obvious thing is that most of these circles are largely red. But also the second thing is that most of the citations are actually in the acknowledgements with the light blue bars here. So most of the mentions are not actually even in the text and connected to the specific work that the researchers are doing. They're just mentioning mentioning that they're broadly used within the acknowledgements of the paper. The thing that I found striking from this work is that there were very few citations. I was really expecting more, because I do know that researchers are using a lot, these libraries a lot, so I kind of wanted to get like a better picture, like a broader picture of how much are these packages used. It's very hard to find statistics on how much packages are being used, but one of the things I thought about was looking at PyPy and they do have a database with all the downloads over the last, I don't know, 20 years, ever since PyPy has been established in a BigQuery database in Google. So I pulled the number of downloads over time. I didn't do it for NumPy, I didn't want to pay a lot of money, but I did it for a number of astronomical libraries and the blue one here is AstroPy to kind of like compare to the previous plots. So AstroPy had only a couple thousand citations last year, but it had 23 million downloads from PyPy. So if you look at the number of downloads per citation that comes to for these small astrophysics libraries between ten thousand and between a thousand and ten thousand downloads from PyPy per citation, So 10,000 at least downloads, 10,000 downloads to generate one citation, which is like an incredibly low yield in terms of citations. Oops, oh, I meant to show this. So this is the downloads per mention in a paper. So this is incredibly low yield of citations per download, which shows that, you know, the number of citations that researchers are looking at may be severely underestimating the usage of their libraries. So implications for all of this. Mentions of major libraries are increasing, but still lag significantly behind the self-reported use. So the mentions alone are not a good indicator of how much these libraries are being used. Policies are good, but community culture is what really matters, and community culture is what we really need to change in order to improve citations. One citation seems to be generated every 1,000 to 10,000 downloads, which is extremely low yield. And since software development is incredibly impactful but lacks credit, this can impact career paths, funding, and maintenance for a lot of open source libraries. So with all of these implications, let me move to the second part of my talk, which is where I want to talk about how credit is a collaborative process. So credit is not just the citations in the journal. Credit is everything that comes before that. Credit is the work that software developers credit. In order for credit to happen, the software developers, the researchers and or users of the software and the journals that publish need to collaborate with each other in order to improve the process. Specifically I want to focus on three different questions here and offer some recommendations on how we can improve the software citation process. First I'm going to talk about how software developers can ensure that they receive appropriate credit for their work. Second I'm going to talk about how we can users can credit effectively the software that they use, and third, I'm going to talk about how journals and repositories can promote proper software citation. So first, how can software developers ensure that they receive proper credit? In the first place, software developers need to provide clear instruction how they want to be credited. Do they want citation to a paper? Do they want a citation to a DOI? Do they want a PayPal donation? Do they want to receive a fruit basket on their door? You need to specify as a developers how you want to receive credit for the project. Some examples of citing of citation recommended citations that I've included here on the side are NumPy which provides the requested citation including also a BibTeX format you can copy and paste directly into your paper and also Matplotlib at the bottom which also requests that you cite the Hunter paper from 2007. The second recommendation is to assign a persistent digital object identifier or DOI to your software, to your library. These are extremely easy to make. For the last 10 years you can log into Zenodo with your GitHub credentials and produce an archive of your repository which will perpetuity holds the contents of your repository. Finally, also, next you can provide a citation CFF file in your GitHub repository. This is very standard but it's actually kind of very underutilized. There's only about 10,000 of these, I think, in all of GitHub. Citation CFF files don't actually contain digital object identifier or an actual, you know, citable entity, but they provide the information that users need in a tab that can easily be accessed from the GitHub repository page where users can copy and paste the citation that you would like them to use. And finally, you can publish your research, you can publish your software in the research or software paper. I I understand this is a big bar. The screenshot on the side is from the AstroPy paper. This is 43 pages. And this is a really big bar. So maybe you don't want to publish a paper that's 43 pages so that you get some citation that people don't even maybe cite properly. So I want to introduce you to the concept of the Journal of Open Source Software, where I'm an editor, but this is not the only journal that focuses on creating entities related to software that look like papers and can be cited like papers in order to create a digital footprint for software. The papers for JAWS specifically are very short. They're just over 1,000, between 1,000 and 2,000 words. So we want to keep them up to three pages. The reviews are really focused on installation, functionality, performance, and documentation of the software, and example usage and tests. They're not actually reviews as much of the paper as of the software itself, and they ensure consistency of all the packages that are being documented under JAWS. We have seven tracks that cover a number of different research areas, and we have published just over 2,500 papers since 2016. If you're interested in publishing in JAWS, come and chat with me, or you can join us as an editor or as a reviewer. You can go to the website, you can sign yourself with your GitHub repo, GitHub ID, and we'll reach out to you if we get submissions relevant to your area. You can also submit your software to a software citation station where researchers can find how to cite it. Next, how can users effectively credit the software they use? Here's some recommendations. If you're a user, use the preferred citation method specified by the author citing both software DOIs and publications if they're recommended. For example, AstroPy requested you cite all three papers that they have published. To ensure reproducibility, you can also cite the hashtag or version that you use specifically for your project. Format the references as citation, not as a footnote, and ensure the citation is in the reference list, which allows systems online and in journals to track the citation and count it properly. Add the software in the acknowledgements if there's no obvious place elsewhere and acknowledge all the tools that are significantly contributed. There's no such thing as too big to cite. And finally, a few recommendations for journals and repositories to promote software citations. Journals should establish clear policies and develop guidelines that require authors to cite software appropriately. They should educate their stakeholders, provide training and resources to editors and authors. They should implement technical solutions to ensure that their platforms support the inclusion and proper formatting of software citations, and they should also monitor compliance. Three of the different journals that I researched have recommendations on software citation, But none of them require consistently all authors to cite software. As a result, the number of citations is well below the usage. These recommendations have brought broader connections to industry. Transparent credit system and citation systems can be useful for researchers outside of academia as persistent identifiers. Publishing software and papers and documentation can encourage clean and maintainable and well-documented code, and this is all really important for sustainability, making sure that contributions are explicitly acknowledged, and this supports healthier teams and long-term, long-list projects. So with that, I'm going to stop. You can find my talk and the repository that produced the plots that I showed on these QR codes, and I have a number of these stickers. I cite code. If you'd like one, find me tomorrow. I forgot to bring them today. I'm really sorry, but I'll be here tomorrow. Find me and I'll give you a sticker. Thank you.
Speaker 2 [22:25]
So, thanks a lot for this nice talk. We have questions. You can use a slider or since also we are like a very small room, I also can just give you the mic. So, okay, there's one question. We'll start with yours and then the next is yours. Do you know a function which lists all DOI of my SBOM software bill of materials?
Speaker 1 [22:50]
Can I see what this looks like?
Speaker 2 [22:51]
Okay, so SPOM, so Software Bill of Materials. Maybe, Reimer, if you want to shortly explain what you mean with that.
Speaker 3 [23:03]
Yeah, I cited CondaForge, for example. I also have a DOI for my project, but when I look up the log file by PixieLog, for example, I see the whole list of all packages I use.
Speaker 1 [23:16]
use
Speaker 3 [23:18]
And, of course, it would be very good if I would cite them all, but it would be nice if we have a function which can maybe pass a catalogue or something else. And the question was how to do this or what is your advice when someone really wants to do this for the whole software bill of materials.
Speaker 1 [23:43]
Yeah, that's a really good question. How do you find all the citations for everything that your package, your library depends on? That's basically the question, right? So the closest it exists is this website that I showed that has citations in one place for a lot of projects. This probably doesn't capture everything that's out there. It's maintained by two people who are doing this on volunteer effort. So, I'm sure they would appreciate contributions. I think the best thing is for authors to contribute their own work. But that's a very difficult question. The issue that I find is, in a lot of cases, even if I want to do the work myself, the authors of a lot of libraries have not provided anything that is citable for me to cite. And that is a major issue. So if you do the work yourself and you find that half of the packages that your library depends on, don't even have a mention how they should be cited, open an issue on their repository and ask them to add something so that, you know, in the future the next person can cite it.
Speaker 2 [24:53]
Next question.
Speaker 4 [24:55]
Hi, thank you for the nice talk. I guess a lot of this problem is going to be cultural from the journal's side. How would you recommend people go about dealing with things like citation limits? Let's say I'm writing a paper and I can only cite 30 things or 50 things, and I, of course, want to cite the scientific parts because I have a limited number of things I can use there. How would you go about solving that problem?
Speaker 1 [25:24]
Yeah, I'd love to. What area of research do you work in?
Speaker 4 [25:28]
I'm a climate scientist so as soon as you start to look at stuff like nature or science they say okay you can have 30 papers and that's it and then you need to make sure that what you cite is relevant to the scientific part of what you're doing. Of course I would also love to cite the path unpackers that I use but it's simply not enough room for them sometimes.
Speaker 1 [25:51]
Yeah, this is a really good point. So astronomy journals generally don't have citation limits But this is something that I think should be a conversation with the journals themselves because software is a really important infrastructure and Yes, research is also important to be cited but not giving credit to software is Yeah, it's really bad So I think that should be on part of the journals to allow citations to software to not count towards the limit of citations
Speaker 2 [26:24]
I also have a question. Your field of research is astronomy and it's kind of very special in a way that it's like a science but doesn't do experiments so it's like more about like viewing what's going on but you can't just blow up some planets and then see how the dust settles and so on.
Speaker 1 [26:40]
and so on.
Speaker 2 [26:41]
And I wonder if your finding is specific for astronomy, or do you see some universal findings regarding other sciences? And also within astronomy, we saw that it's like from 1% to 12%, kind of a high difference between these journals. So what's your guess about all sciences, how do you have some?
Speaker 1 [27:08]
No, I don't, that's a very good question, and that's one of the reasons I actually did this work, and that was my motivation, was to understand how things vary within astronomy, but also to hopefully follow up or collaborate with others on exploring other fields. So if you're in a field where you'd like to do this kind of work, I'd be happy to work with you, or if you have suggestions for other analysis that you'd like to see on this data, I'd be happy to chat with you. But yeah, I'm happy to take recommendations. Feel free to find me on GitHub, Blue Sky, whatever. Feel free to open an issue on the repository. It's a public repository. So I'd love to hear feedback on this. I literally finished this yesterday.
Speaker 2 [27:52]
Other questions Okay, so so yeah
Speaker 4 [27:59]
So my question is, so with software having multiple versions and if we have DOIs for each of them, it would be very difficult to track them under a common name. So what can be a solution, do you think?
Speaker 1 [28:14]
Yeah, so Zenodo actually allows you to have a DOI for all the versions that refers to the package itself independent of the version that you're citing specifically. Or you can cite the DOI for the specific version that you're using. So that's a functionality that exists on Zenodo for the last few years. So that's definitely something you can do.
Speaker 4 [28:40]
With papers, it would be a bit more difficult because once it's published, it's more or less fixed.
Speaker 1 [28:46]
Yes, yeah, so that's definitely an issue, but for the paper you can also link, so for JAWS, for example, Journal of Open Source Software, we ask the authors to make another repository and we link to that, so if you're citing the paper on JAWS, you can also cite the UI for the whole package.
Speaker 2 [29:14]
Thanks again, Eva, for this nice talk.