Random Similarity Forests
Authors:
Maciej Piernik,
Dariusz Brzezinski,
Pawel Zawadzki
Abstract:
The wealth of data being gathered about humans and their surroundings drives new machine learning applications in various fields. Consequently, more and more often, classifiers are trained using not only numerical data but also complex data objects. For example, multi-omics analyses attempt to combine numerical descriptions with distributions, time series data, discrete sequences, and graphs. Such…
▽ More
The wealth of data being gathered about humans and their surroundings drives new machine learning applications in various fields. Consequently, more and more often, classifiers are trained using not only numerical data but also complex data objects. For example, multi-omics analyses attempt to combine numerical descriptions with distributions, time series data, discrete sequences, and graphs. Such integration of data from different domains requires either omitting some of the data, creating separate models for different formats, or simplifying some of the data to adhere to a shared scale and format, all of which can hinder predictive performance. In this paper, we propose a classification method capable of handling datasets with features of arbitrary data types while retaining each feature's characteristic. The proposed algorithm, called Random Similarity Forest, uses multiple domain-specific distance measures to combine the predictive performance of Random Forests with the flexibility of Similarity Forests. We show that Random Similarity Forests are on par with Random Forests on numerical data and outperform them on datasets from complex or mixed data domains. Our results highlight the applicability of Random Similarity Forests to noisy, multi-source datasets that are becoming ubiquitous in high-impact life science projects.
△ Less
Submitted 11 April, 2022;
originally announced April 2022.
Single-molecule imaging of DNA gyrase activity in living Escherichia coli
Authors:
Mathew Stracy,
Adam J. M. Wollman,
Elzbieta Kaja,
Jacek Gapinski,
Ji-Eun Lee,
Victoria A. Leek,
Shannon J. McKie,
Lesley A. Mitchenall,
Anthony Maxwell,
David J. Sherratt,
Mark C. Leake,
Pawel Zawadzki
Abstract:
Bacterial DNA gyrase introduces negative supercoils into chromosomal DNA and relaxes positive supercoils introduced by replication and transiently by transcription. Removal of these positive supercoils is essential for replication fork progression and for the overall unlinking of the two duplex DNA strands, as well as for ongoing transcription. To address how gyrase copes with these topological ch…
▽ More
Bacterial DNA gyrase introduces negative supercoils into chromosomal DNA and relaxes positive supercoils introduced by replication and transiently by transcription. Removal of these positive supercoils is essential for replication fork progression and for the overall unlinking of the two duplex DNA strands, as well as for ongoing transcription. To address how gyrase copes with these topological challenges, we used high-speed single-molecule fluorescence imaging in live Escherichia coli cells. We demonstrate that at least 300 gyrase molecules are stably bound to the chromosome at any time, with ~12 enzymes enriched near each replication fork. Trapping of reaction intermediates with ciprofloxacin revealed complexes undergoing catalysis. Dwell times of ~2 s were observed for the dispersed gyrase molecules, which we propose maintain steady-state levels of negative supercoiling of the chromosome. In contrast, the dwell time of replisome-proximal molecules was ~8 s, consistent with these catalyzing processive positive supercoil relaxation in front of the progressing replisome.
△ Less
Submitted 2 November, 2018; v1 submitted 1 November, 2018;
originally announced November 2018.