Tuesday, April 11, 2023

Recommendation systems for papers

The volume of papers being published and preprints being submitted to the arXiv has grown enormously since the time when I started my PhD studies. Moreover, competing preprint platforms such as Optica Open have been launched. This deluge of papers is impossible to keep up with. 

Therefore, there is growing interest in developing advanced bibliometric tools in order to stay abreast of the most important developments in one's own research field. There are two distinct approaches to solving this issue. 

The first is crowdsourcing. Websites such as scirate (founded by and most widely used by quantum physicists) and PubPeer (seems to be popular in life sciences) and allow users to upvote and comment on preprints they think are interesting. Preprints that are upvoted more appear higher on the page and therefore receive more attention and more views. The idea then is that the most important works will be upvoted more and will be seen by more people.

Other approaches are based on machine learning or bibliometric analysis methods where a new paper is analyzed by model that takes as its input various attributes of the paper, such as the authors, the topic, keywords, and its reference list. Models can be trained to pick out papers which are likely to relevant or more important and show them to the end user. 

One example of this now supported by arXiv is Litmaps, which constructs a graph that places an article in the context of previous and subsequent works. The visualisation can be customised to highlight different features, such as the publication date and total number of citations in the plot below. It seems at least that for the example below the "Seed" map seems biased towards highlighting review articles (missing hot recent results such as those on quantized Thouless pumping of solitons). The other visualisations offered are "Discover" (for finding overlooked papers) and "Map" (for telling a "research story"), but they require an account, and presumably a subscription for serious use.

Litmap of "Edge solitons in nonlinear photonic topological insulators"

Each approach has its own pros and cons. One problem with the crowdsourcing approach is that it can be sensitive to initial perturbations and tends to amplify existing well-known authors while inhibiting the promotion of less well-known authors. For example, people may upvote a paper just because it is written by familiar names and the title and abstract look interesting, leading to a winner-takes-all effect. 

This I think is a particularly important problem in research. At least my style is that I prefer not to work on ideas that are already very popular. I think to really make a deep breakthrough in research we need to see something that nobody else has noticed before. This often will involve getting insights from papers which have been forgotten or overlooked by the wider community. In this context crowdsourcing has the danger of leading to a groupthink in which the popularity of certain topics may exceed their promise, due to people working on them simply because many other people are also working on them. Thus, the choices of a few early adopters or "academic influencers" will end up getting amplified more and more.

The machine learning approach has the potential to give a more thorough and systematic coverage of the literature by being able to analyze all new papers and find the most interesting ones without being subject to this reliance on or sensitivity to initial fluctuations and early upwards. While there is promise of course, the big question is how are you supposed to develop a model to rank individual papers without studying deeply the science they contain? And how much can you trust a proprietary, closed-source model whose inner workings and potential biases are unknown?

For example, a crude first approximation might involve analyzing the references of a new paper to see what is cited. If a new paper cites important previous results (which can be estimated by how often they have been cited) then hopefully the paper will be worth reading. However simply counting the raw citations or references will be prone to bias. Different fields have different standards as to what is and should be cited. In some fields now you will see paper introductions which cite dozens or even hundreds of papers. In this case the value of an individual citation is relatively low and so looking at the citations alone we won't give much information as to what the paper is roughly about or whether it is worth reading.

It seems therefore that more sophisticated approaches are necessary. The context of a citation is important. Papers cited in bulk are not that valuable. For example, in an introductory paragraph "x has received a lot of interest lately [1-103]" only reveals that x is a hot topic. On the other hand, a sentence like "We apply the method of Ref. [16]..." tell us that there is probably a very close connection to whatever Ref. [16] is about. Thus, the integration of large language models such as BioGPT with paper recommendation systems is likely to improve their performance, thereby greatly improving the productivity of researchers who use them.

What do you think? Are there any other tricks for keeping up with the literature in your field?

Thursday, April 6, 2023

PhD position in Floquet topological photonics (experimental)

The group of Alberto Amo at PhLAM laboratory (CNRS / University of Lille) has a PhD position opening in Floquet topological photonics (experimental). The PhD thesis is part of the ERC EmergenTopo project and aims at studying linear and nonlinear topological phases in time modulated synthetic lattices. More information about the project can be found in the application website.

According to the job advertisement, the PhD project will broadly study fractal lattices and photon-photon interaction effects in coupled optical fibers, building on the group's recent work published in Physical Review Letters. This work falls within the emerging research direction of synthetic dimensions in photonics, which uses spatial or periodic modulations to small systems of a few coupled resonators or waveguides to emulate higher-dimensional lattices. There are already several reviews on this topic, including "Synthetic dimension in photonics", "Topological quantum matter in synthetic dimensions", and "Topological photonics in synthetic dimensions".

 
Application deadline: 5th May 2023.

Starting date: 1st October 2023



Tuesday, April 4, 2023

Conferences in Ukraine

The first two international physics conferences I had the pleasure of attending, way back in September 2011, were held in Ukraine. This was also my first time travelling outside Australia. 

After 30 hours and 4 flights I landed in Kharkiv for the first conference, the International Workshop on Nonlinear Photonics (NLP*2011), held at the Kharkiv National University, located at the picturesque Svobody Square in the city centre.


Svobody Square. The statue of Lenin was torn down in 2014, replaced by a fountain in 2020, and presumably shelled in 2022.

Kharkiv was the home and origin of many great theoretical physicists. Landau and Lifschitz began writing their classic textbook series Course of Theoretical Physics there. Outside the auditorium in which the conference was held, attendees were greeted by an impressive Soviet-era mural.

The entrance to the conference auditorium. The university buildings were destroyed by Russian army shelling in March 2022.

This workshop was my first chance to meet many leading researchers working on nonlinear optics and singular optics, including the late Marat Soskin, who gave a memorable talk on the creation and destruction of topological defects in nematic liquid crystals. One afternoon my then-colleague and future Ignobel Prize Laureate Ivan Maksymov showed me around the city, which was where he had completed his physics studies.

 The following week I attended the Tenth International Conference on Correlation Optics, held at Chernivsti National University on the other side of the country, two flights and a train ride away.

 

The beautiful grounds of the Chernivsti National University, a UNESCO World Heritage Site, constructed between 1864 and 1882.

The Correlation Optics conference series is still going strong, with the next edition planned to be held in hybrid mode in September 2023. As one attendee aptly put it, "No one really knows what correlation optics is precisely, so its themes can continuously adapt to changes in research trends." At the time, one trend was the increasing availability of nanofabrication facilities leading to a transition from micro-optics to nano-optics.

Before dawn on the morning after the end of the conference, what seemed like all of the international attendees converged on the tiny Chernivsti Airport to catch the only international flight running that day. After "checking in" our baggage, we had to wheel it ourselves onto the tarmac to be loaded onto the small jet plane while we considered holding a post-conference session during the flight. 

I travelled onwards to Germany to visit collaborators at the University of Münster. Due to a chance encounter at ICOAM last year, we resumed our collaboration leading to a paper soon to be published in Photonics Research. But that's another story.

Thursday, March 30, 2023

arXiv highlights

Here are some papers that caught my eye over the past month:


Germain Curvature: The Case for Naming the Mean Curvature of a Surface after Sophie Germain

This essay argues that the intrinsic curvature of a surface, aka the Gaussian curvature, should be named instead the Germain curvature, since Gauss was not the first to study it.

I remember attending a lecture by Sir Michael Berry (of Berry phase fame) where he made a compelling argument against naming new objects or effects after people, on account of the three "Laws of Discovery":

"1. Discoveries are rarely attributed to the correct person

2.Nothing is ever discovered for the first time

3. To come near to a true theory, and to grasp its precise application, are two very different things, as the history of science teaches us. Everything of importance has been said before by someone who did not discover it."

Indeed, versions of the Berry phase had been previously decades before Berry, by Pancharatnam, Rytov, and others. For this reason he prefers the name "geometric phase." Similarly, intrinsic curvature is perhaps a more suitable alternative to Gaussian curvature.

The problem with naming effects after people is that the nature of the effect becomes opaque unless one already knows what it means. The situation becomes even worse when different groups decide to name the same effect after different people. On the other hand, simple yet descriptive names including geometric phase and intrinsic curvature reveal some sense of what is meant to the outsider. The absence of a simple-sounding name may indicate that we don't really understand the effect.

An Aperiodic Monotile

The authors discover a family of shapes that can tile the 2D plane, but only aperiodically. The shapes are non-convex mirror-asymmetric polygons. Tiling the plane involves placing a mixture of the polygon and its reflection, but the two can never be arranged to form a regular pattern. Can this kind of aperiodic tiling lead to novel physical properties of some system or model? For example, tight binding lattices can be obtained from tilings by identifying corners as "sites", with coupling between sites linked by edges of the tiling shapes.

Spectral localizer for line-gapped non-Hermitian systems

The localizer theory I have discussed previously (here and here) is now generalized to non-Hermitian systems! This is relevant to understanding the properties and robustness of certain topological laser models.

A quantum spectral method for simulating stochastic processes, with applications to Monte Carlo

This preprint shows that the quantum Fourier transform can be used to efficiently simulate random processes such as Brownian motion. In contrast to previous "digital" quantum Monte-Carlo approaches, here the authors consider an encoding in which the value of the random variable is encoded in the amplitude of the quantum state, with different basis vectors corresponding to different time steps. Since Prakash's earlier work on quantum machine learning using subspace states was the inspiration of our recent quantum chemistry work I think this paper is well worth a closer read!

Photonic quantum computing with probabilistic single photon sources but
without coherent switches

 If you want to learn more about the photonic approach for building a fault tolerant quantum computer (being pursued by PsiQ), you should read Terry Rudolph's always-entertaining papers. Even though the approaches presented in this manuscript (first written in 2016-2018) are now obsolete this is still well worth a read as a resource on the key ingredients of potentially-scalable methods for linear optical quantum computing.

An Improved Classical Singular Value Transformation for Quantum Machine Learning

The field of quantum machine learning has seen two phases. The first phase was sparked by the discovery of the HHL algorithm. HHL and its descendants promised an exponential speedup for certain linear algebra operations appearing widely-used machine learning techniques, arguably triggering the current boom in quantum technologies. However, running these algorithms on any useful problem will require a full fault-tolerant quantum computer.

Consequently, novel quantum algorithms for machine learning have attracted interest as a possible setting for achieving useful quantum speedups before a large scale fault-tolerant quantum computer can be developed. The power of these newer algorithms is much less certain and still under intense debate. Nevertheless, researchers could find solace in the hope that, even if these NISQ-friendly algorithms do not end up being useful, eventually we will achieve a quantum advantage using HHL-based algorithms.

The dequantization techniques pioneered by Ewin Tang and collaborators are starting to suggest that even a quantum advantage based on fault-tolerant algorithms such as HHL may turn out to be a mirage. This paper presents a new efficient classical sampling-based algorithm that reduces the potential quantum speedup for singular value transformations from exponential to polynomial. This affects a variety of quantum machine learning algorithms, including those for topological data analysis, recommendation systems, and linear regression.

 

 

Tuesday, March 28, 2023

TDA Week 2023

TDA Week is a five-day conference on topics related to topological data analysis, held this year in person at Kyoto University from July 31 (Mon) to August 4 (Fri). It follows last year's conference, which was held online due to covid restrictions.

The abstract submission deadline for poster presentations is 14th April. Presenting authors may request partial support for travel expenses.

Wednesday, March 22, 2023

Predatory publishing and open access

I recently stumbled upon Predatory Reports, an anonymously-run website that lists journals and publishers with dubious practices and standards. This is a growing problem with the rise of open access publishing mandates; since authors only pay if their article is published, there is an incentive to lower standards and publish everything.

It is interesting to note the inclusion of MDPI and Frontiers Media in the Predatory Reports list. All of the justifying examples are, to the best of my knowledge, taken from life sciences journals, and it is not clear whether similar issues affect their physics journals. Personally, however, I have received occasional review requests from them for papers which I clearly have no expertise in reviewing. 

A bigger issue (particularly with MDPI) is their spamming of special issue invitations. Since the guest editors will nominally handle submissions, including selecting potential referees, this can lead large variations in quality and standards among the articles published in a particular journal. Paolo Crosetto has a blog post analysing of the business model of special issue publishing and how it has turned into a money-printing machine for MDPI.

In related news, Nature published a feature on the journal eLife's decision last year to switch to a "publish everything" model, in which all papers which are sent to peer review are published alongside the referee reports. Nature is itself experimenting with similar open review ideas and the potential for the role of journals to shift from selective publishing to obtaining credible peer review reports. This model is particularly attractive for for-profit publishers, since it offers an attractive and reliable new source of revenue - under the open access model a journal loses money on every paper it rejects.

What will probably limit uptake of the publish everything model is that authors are ultimately after visibility of their work. Visibility requires selectivity, and you cannot have selectivity without rejecting a lot of papers.