OPENADMET CYP INHIBITION CHALLENGE · METHOD REPORT

FreeT4-CYP

Molecular fingerprints, tree models, and masked multi-task learning for CYP inhibition. New cloud charges: $0.00.

Report generated 2026-09-26 05:36 UTC. This describes our method and internal validation. It does not claim a blind-test leaderboard score.

Data and provenance

We use the public OpenADMET training data: direct inhibition, TDI labels, and the single-concentration screening table. Dataset discovery and schema inspection use Scigantic’s public MCP. No proprietary or external training data are used.

There are 6,146 normalized training molecules and 750 blinded test rows in the prepared experiment. Original test identifiers and SMILES are preserved. Local artifacts record dataset revisions, checksums, dependencies, and molecule split membership.

Representations and models

SMILES are normalized to molecular parents while preserving stereochemistry in molecular identity. Morgan count fingerprints (radius 2, 2,048 dimensions, log1p, chirality disabled) are combined with RDKit 2D descriptors. Descriptor imputation and scaling are fitted on each training partition only.

The baseline fits separate LightGBM models to four pIC50 endpoints and two TDI labels: 600 trees, learning rate 0.03, 31 leaves, L2 regularization 1. The neural alternatives share dense layers of 512 and 256 units with ReLU and dropout 0.15, trained using AdamW (learning rate 0.001, weight decay 0.0001), batches of 256, and at most 60 epochs.

Regression uses masked distance to the measured credible interval; classification uses masked binary cross-entropy. One neural variant also predicts four primary-screen log2 fold changes as auxiliary targets, using standardized MSE. Per-task means prevent dense endpoints dominating; family weights are 1 for regression, 1 for classification, and 0.25 for auxiliary targets. Missing labels are not treated as negatives. Screening measurements are never inference inputs.

Model selection

A fixed scaffold-group split holds out 1,244 molecules, leaving 4,902 for three grouped development folds. Duplicate structures and their auxiliary measurements remain together. Molecules without a Murcko scaffold are grouped by parent connectivity. Seed: 42.

Each track is selected independently using development predictions. TDI thresholds maximize development MCC over 0.05–0.95 in increments of 0.01, with ties resolved toward 0.5. The selected regression model is lgbm; classification uses mlp with CYP2D6/CYP3A4 thresholds [0.21, 0.21]. Choices and thresholds are frozen before evaluating the holdout.

CandidateDevelopment MA-ST-RAE ↓Development macro MCC ↑
lgbm0.74940.2438
mlp1.75180.2581
mlp_aux1.70850.2481

Development metrics include early-stopping and threshold selection and are not unbiased estimates. Final models are refitted on all training molecules using the median selected development epoch count for neural models.

Untouched holdout evaluation

Macro ST-RAE: 0.7070. Macro MCC: 0.2159.

EndpointLabeled moleculesMetricScore95% bootstrap interval
CYP1A2_pIC50_direct_inhibition264st_rae0.81820.7222–0.9314
CYP2C9_pIC50_direct_inhibition245st_rae0.63420.5401–0.7492
CYP2D6_pIC50_direct_inhibition316st_rae0.86930.8239–0.9242
CYP3A4_pIC50_direct_inhibition437st_rae0.50640.4402–0.5740
CYP2D6_is_TDI316mcc0.0964-0.0213–0.2281
CYP3A4_is_TDI730mcc0.33530.2615–0.4151

Intervals use 1,000 paired molecule-row bootstrap resamples with fixed thresholds. ST-RAE applies credible-interval distance to both model error and the mean-predictor denominator, matching the pinned official scoring code.

Compute and reproducibility

Executed model backends: cpu, mps. CPU/MPS denotes local Apple hardware; CUDA denotes an NVIDIA runtime. Scigantic advertises a free, preemptible T4 tier; availability and use must not be inferred from this entry’s alias. Actual paid cloud resources used: none. New service charges: $0.00.

Epoch checkpoints retain model, optimizer, preprocessing, and random states. Source, dependency lock, checkpoints, split manifests, development predictions, and validation results are retained locally. Training source and model weights are not publicly released.

Limitations and submission checks

The training matrix is sparse and selected by assay screening, while the blind test is a dense analog-expansion set. Scaffold holdout performance is not a guarantee of blind-test performance. Row-bootstrap intervals do not capture all dependence within chemical families or the full uncertainty of model selection.

We submit two independent 750-row Parquet files. Checks enforce exact original identifiers/SMILES, finite regression outputs, sample standard deviation at least 0.01 per regression endpoint, boolean classification outputs with both classes, and official schema validation. We do not alter predictions artificially to pass variability checks or tune models on leaderboard feedback.