xgboost

Author	SHA1	Message	Date
Jiaming Yuan	e1a2c1bbb3	[EM] Merge GPU partitioning with histogram building. (#10766 ) - Stop concatenating pages if there's no subsampling. - Use a single iteration for histogram build and partitioning.	2024-08-31 03:25:37 +08:00
Jiaming Yuan	61dd854a52	[EM] Refactor GPU histogram builder. (#10764 ) - Expose the maximum number of cached nodes to be consistent with the CPU implementation. Also easier for testing. - Extract the subtraction trick for easier testing. - Split up the `GradientQuantiser` to avoid circular dependency.	2024-08-30 02:39:14 +08:00
Jiaming Yuan	4fe67f10b4	[EM] Have one partitioner for each batch. (#10760 ) - Initialize one partitioner for each batch. - Collect partition size during initialization. - Support base ridx in the finalization.	2024-08-29 01:35:17 +08:00
Jiaming Yuan	d6ebcfb032	[EM] Support CPU quantile objective for external memory. (#10751 )	2024-08-27 04:16:57 +08:00
Jiaming Yuan	142bdc73ec	[EM] Support SHAP contribution with QDM. (#10724 ) - Add GPU support. - Add external memory support. - Update the GPU tree shap.	2024-08-22 05:25:10 +08:00
Jiaming Yuan	2258bc870d	Add more tests and doc for QDM. (#10692 )	2024-08-16 23:30:04 +08:00
Jiaming Yuan	827d0e8edb	[breaking] Bump Python requirement to 3.10. (#10434 ) - Bump the Python requirement. - Fix type hints. - Use loky to avoid deadlock. - Workaround cupy-numpy compatibility issue on Windows caused by the `safe` casting rule. - Simplify the repartitioning logic to avoid dask errors.	2024-07-30 17:31:06 +08:00
Jiaming Yuan	e8a962575a	[EM] Allow staging ellpack on host for GPU external memory. (#10488 ) - New parameter `on_host`. - Abstract format creation and stream creation into policy classes.	2024-06-28 04:42:18 +08:00
Jiaming Yuan	73afef1a6e	Fixes for numpy 2.0. (#10252 )	2024-05-07 03:54:32 +08:00
github-actions[bot]	2925cebdca	[CI] Use latest RAPIDS; Pandas 2.0 compatibility fix (#10175 ) * [CI] Update RAPIDS to latest stable * [CI] Use rapidsai stable channel; fix syntax errors in Dockerfile.gpu * Don't combine astype() with loc() * Work around https://github.com/dmlc/xgboost/issues/10181 * Fix formatting * Fix test --------- Co-authored-by: hcho3 <hcho3@users.noreply.github.com> Co-authored-by: Hyunsu Cho <chohyu01@cs.washington.edu>	2024-04-15 13:38:53 -07:00
Jiaming Yuan	e14c3b9325	Optional normalization for learning to rank. (#10094 )	2024-03-08 12:41:21 +08:00
Jiaming Yuan	69a17d5114	Fix with None input. (#10052 )	2024-02-20 22:34:22 +08:00
Jiaming Yuan	54b71c8fba	Fix with black 24.1.1. (#10014 )	2024-01-30 17:24:11 +08:00
Philip Hyunsu Cho	c8f5d190c6	[CI] Stop Windows pipeline upon a failing pytest (#10003 )	2024-01-24 22:54:21 -08:00
Jiaming Yuan	d12cc1090a	Refactor tests for training continuation. (#9997 )	2024-01-24 16:07:19 +08:00
Jiaming Yuan	9f73127a23	Cleanup Python GPU tests. (#9934 ) * Cleanup Python GPU tests. - Remove the use of `gpu_hist` and `gpu_id` in cudf/cupy tests. - Move base margin test into the testing directory.	2024-01-04 13:15:18 +08:00
Jiaming Yuan	c3a0622b49	Fix using categorical data with the score function of ranker. (#9753 )	2023-11-07 07:29:11 +08:00
Jiaming Yuan	adea842c83	Fix inplace predict with fallback when base margin is used. (#9536 ) - Copy meta info from proxy DMatrix. - Use `std::call_once` to emit less warnings.	2023-09-05 01:04:24 +08:00
Jiaming Yuan	972730cde0	Use matrix for gradient. (#9508 ) - Use the `linalg::Matrix` for storing gradients. - New API for the custom objective. - Custom objective for multi-class/multi-target is now required to return the correct shape. - Custom objective for Python can accept arrays with any strides. (row-major, column-major)	2023-08-24 05:29:52 +08:00
Jiaming Yuan	19b59938b7	Convert input to str for hypothesis note. (#9480 )	2023-08-15 02:27:58 +08:00
Jiaming Yuan	54029a59af	Bound the size of the histogram cache. (#9440 ) - A new histogram collection with a limit in size. - Unify histogram building logic between hist, multi-hist, and approx.	2023-08-08 03:21:26 +08:00
Jiaming Yuan	912e341d57	Initial GPU support for the approx tree method. (#9414 )	2023-07-31 15:50:28 +08:00
Jiaming Yuan	275da176ba	Document for device ordinal. (#9398 ) - Rewrite GPU demos. notebook is converted to script to avoid committing additional png plots. - Add GPU demos into the sphinx gallery. - Add RMM demos into the sphinx gallery. - Test for firing threads with different device ordinals.	2023-07-22 15:26:29 +08:00
Jiaming Yuan	0a07900b9f	Fix integer overflow. (#9380 )	2023-07-15 21:11:02 +08:00
Jiaming Yuan	9da5050643	Turn warning messages into Python warnings. (#9387 )	2023-07-15 07:46:43 +08:00
Jiaming Yuan	04aff3af8e	Define the new `device` parameter. (#9362 )	2023-07-13 19:30:25 +08:00
Jiaming Yuan	20c52f07d2	Support exporting cut values (#9356 )	2023-07-08 15:32:41 +08:00
Jiaming Yuan	39390cc2ee	[breaking] Remove the `predictor` param, allow fallback to prediction using `DMatrix`. (#9129 ) - A `DeviceOrd` struct is implemented to indicate the device. It will eventually replace the `gpu_id` parameter. - The `predictor` parameter is removed. - Fallback to `DMatrix` when `inplace_predict` is not available. - The heuristic for choosing a predictor is only used during training.	2023-07-03 19:23:54 +08:00
Jiaming Yuan	ee6809e642	Use mmap for external memory. (#9282 ) - Have basic infrastructure for mmap. - Release file write handle.	2023-06-19 18:52:55 +08:00
Jiaming Yuan	9fbde21e9d	Rework the precision metric. (#9222 ) - Rework the precision metric for both CPU and GPU. - Mention it in the document. - Cleanup old support code for GPU ranking metric. - Deterministic GPU implementation. * Drop support for classification. * type. * use batch shape. * lint. * cpu build. * cpu build. * lint. * Tests. * Fix. * Cleanup error message.	2023-06-02 20:49:43 +08:00
Jiaming Yuan	3913ff470f	Import data lazily during tests. (#9176 )	2023-05-23 03:58:31 +08:00
Jiaming Yuan	08ce495b5d	Use Booster context in DMatrix. (#8896 ) - Pass context from booster to DMatrix. - Use context instead of integer for `n_threads`. - Check the consistency configuration for `max_bin`. - Test for all combinations of initialization options.	2023-04-28 21:47:14 +08:00
Jiaming Yuan	ef13dd31b1	Rework the NDCG objective. (#9015 )	2023-04-18 21:16:06 +08:00
Jiaming Yuan	151882dd26	Initial support for multi-target tree. (#8616 ) * Implement multi-target for hist. - Add new hist tree builder. - Move data fetchers for tests. - Dispatch function calls in gbm base on the tree type.	2023-03-22 23:49:56 +08:00
Jiaming Yuan	5891f752c8	Rework the MAP metric. (#8931 ) - The new implementation is more strict as only binary labels are accepted. The previous implementation converts values greater than 1 to 1. - Deterministic GPU. (no atomic add). - Fix top-k handling. - Precise definition of MAP. (There are other variants on how to handle top-k). - Refactor GPU ranking tests.	2023-03-22 17:45:20 +08:00
Jiaming Yuan	f186c87cf9	Check inf in data for all types of DMatrix. (#8911 )	2023-03-15 11:24:35 +08:00
Jiaming Yuan	7eba285a1e	Support sklearn cross validation for ranker. (#8859 ) * Support sklearn cross validation for ranker. - Add a convention for X to include a special `qid` column. sklearn utilities consider only `X`, `y` and `sample_weight` for supervised learning algorithms, but we need an additional qid array for ranking. It's important to be able to support the cross validation function in sklearn since all other tuning functions like grid search are based on cross validation.	2023-03-07 00:22:08 +08:00
Jiaming Yuan	228a46e8ad	Support learning rate for zero-hessian objectives. (#8866 )	2023-03-06 20:33:28 +08:00
Jiaming Yuan	6a892ce281	Specify src path for isort. (#8867 )	2023-03-06 17:30:27 +08:00
Rory Mitchell	69a50248b7	Fix scope of feature set pointers (#8850 ) --------- Co-authored-by: Jiaming Yuan <jm.yuan@outlook.com>	2023-03-02 12:37:14 +08:00
Philip Hyunsu Cho	6d8afb2218	[CI] Require C++17 + CMake 3.18; Use CUDA 11.8 in CI (#8853 ) * Update to C++17 * Turn off unity build * Update CMake to 3.18 * Use MSVC 2022 + CUDA 11.8 * Re-create stack for worker images * Allocate more disk space for Windows * Tempiorarily disable clang-tidy * RAPIDS now requires Python 3.10+ * Unpin cuda-python * Use latest NCCL * Use Ubuntu 20.04 in RMM image * Mark failing mgpu test as xfail	2023-03-01 09:22:24 -08:00
Jiaming Yuan	cce4af4acf	Initial support for quantile loss. (#8750 ) - Add support for Python. - Add objective.	2023-02-16 02:30:18 +08:00
Jiaming Yuan	457f704e3d	Add quantile metric. (#8761 )	2023-02-13 19:07:40 +08:00
Jiaming Yuan	8a16944664	Fix ranking with quantile dmatrix and group weight. (#8762 )	2023-02-10 20:32:35 +08:00
Rory Mitchell	7214a45e83	Fix different number of features in gpu_hist evaluator. (#8754 )	2023-02-06 23:15:16 +08:00
Jiaming Yuan	0e61ba57d6	Fix GPU L1 error. (#8749 )	2023-02-04 03:02:00 +08:00
Jiaming Yuan	d6018eb4b9	Remove all use of `DeviceQuantileDMatrix`. (#8665 )	2023-01-17 00:04:10 +08:00
Jiaming Yuan	badeff1d74	Init estimation for regression. (#8272 )	2023-01-11 02:04:56 +08:00
Jiaming Yuan	40343c8ee1	Test dask demos. (#8557 ) Co-authored-by: Philip Hyunsu Cho <chohyu01@cs.washington.edu>	2022-12-13 18:37:31 +08:00
Jiaming Yuan	157e98edf7	Support half type from cupy. (#8487 )	2022-11-30 17:56:42 +08:00

1 2 3 4 5 ...

254 Commits