Logo image
Large-scale information retrieval and correction of noisy pharmacogenomic datasets through residual thresholded deep matrix factorization
期刊文章   開放取用(OA)   同儕審查

Large-scale information retrieval and correction of noisy pharmacogenomic datasets through residual thresholded deep matrix factorization

Zhiyue Tom Hu, Yaodong Yu, Ruoqiao Chen, Shan-Ju Yeh, Bin Chen 和 Haiyan Huang
Briefings in bioinformatics, 卷.26(3)
27/05/2025
PMID: 40420482
Web of Science ID: WOS:001495440400001

摘要

Problem Solving Protocol pharmacogenomics datasets drug sensitivity data deep matrix factorization noisy data open-sourced
Pharmacogenomics studies are attracting an increasing amount of interest from researchers in precision medicine. The advances in high-throughput experiments and multiplexed approaches allow the large-scale quantification of drug sensitivities in molecularly characterized cancer cell lines (CCLs), resulting in a number of open drug sensitivity datasets for drug biomarker discovery. However, a significant inconsistency in drug sensitivity values among these datasets has been noted. Such inconsistency indicates the presence of substantial noise, subsequently hindering downstream analyses. To address the noise in drug sensitivity data, we introduce a robust and scalable deep learning framework, Residual Thresholded Deep Matrix Factorization (RT-DMF). This method takes a single drug sensitivity data matrix as its sole input and outputs a corrected and imputed matrix. Deep matrix factorization (DMF) excels at uncovering subtle patterns, due to its minimal reliance on data structure assumptions. This attribute significantly boosts DMF’s ability to identify complex hidden patterns among nuisance effects in the data, thereby facilitating the detection of signals that are therapeutically relevant. Furthermore, RT-DMF incorporates an iterative residual thresholding procedure, which plays a crucial role in retaining signals more likely to hold therapeutic importance. Validation using simulated datasets and real pharmacogenomics datasets demonstrates the effectiveness of our approach in correcting noise and imputing missing data in drug sensitivity datasets (open-source package available at https://github.com/tomwhoooo/rtdmf).

檔案與連結 (2)

pdf
bbaf226 (1)2.91 MB下載檢視
開放存取(OA)
url
https://doi.org/10.1093/bib/bbaf226檢視
已出版(紀錄版本)

相關連結

聯合國永續發展目標(SDGs)

此研究成果有助於達成以下目標:

#3 Good Health and Well-Being

來源:SDGs的產出

詳細資料

Logo image