Commit Graph
638 Commits
Author SHA1 Message Date
Hendrik Groove 9659d0e7bd WARP_SIZE 2024-10-21 21:35:25 +02:00
Hendrik Groove 94ffd57641 try direct instantiation of EvaluateSplitsKernel 2024-10-21 21:22:57 +02:00
Hendrik Groove 8c15f3b665 revert 2024-10-21 11:43:29 +02:00
Hendrik Groove bb2feab0b2 try 2024-10-21 01:55:41 +02:00
Hui Liu b27f35e270 rm hip from src 2024-04-22 12:31:14 -07:00
Hui Liu 8b75204fed merge latest change from upstream 2024-04-22 09:35:31 -07:00
Jiaming Yuan 3f64b4fde3 [coll] Add global functions. (#10203) 2024-04-19 03:17:23 +08:00
Hui Liu ff549ae933 sync upstream code 2024-03-20 16:14:38 -07:00
Jiaming Yuan 53fc17578f Use std::uint64_t for row index. (#10120)
- Use std::uint64_t instead of size_t to avoid implementation-defined type.
- Rename to bst_idx_t, to account for other types of indexing.
- Small cleanup to the base header.
2024-03-15 18:43:49 +08:00
Hui Liu 968dbf25fb merge latest changes 2024-03-12 09:13:09 -07:00
Jiaming Yuan 2c13f90384 Support graphviz plot for multi-target tree. (#10093) 2024-03-09 05:35:25 +08:00
Jiaming Yuan 3941b31ade Disable column sample by node for the exact tree method. (#10083) 2024-03-01 14:16:10 +08:00
Jiaming Yuan 5ac233280e Require context in aggregators. (#10075) 2024-02-28 03:12:42 +08:00
Jiaming Yuan 0ce4372bd4 Use UBJSON for serializing splits for vertical data split. (#10059) 2024-02-25 00:18:23 +08:00
Hui Liu 44db1cef54 Merge branch 'master' into sync-2024Jan24 2024-02-01 14:41:48 -08:00
Philip Hyunsu ChoandJiaming Yuan 4dfbe2a893 [CI] Test building for 32-bit arch (#10021)
* [CI] Test building for 32-bit arch

* Update CMakeLists.txt

* Fix yaml

* Use Debian container

* Remove -Werror for 32-bit

* Revert "Remove -Werror for 32-bit"

This reverts commit c652bc6a037361bcceaf56fb01863210b462793d.

* Don't error for overloaded-virtual warning

* Ignore some warnings from dmlc-core

* Fix compiler warnings

* Fix formatting

* Apply suggestions from code review

Co-authored-by: Jiaming Yuan <jm.yuan@outlook.com>

* Add more cast

---------

Co-authored-by: Jiaming Yuan <jm.yuan@outlook.com>
2024-01-31 13:20:51 -08:00
Hui Liu 3fe874078c merge latest changes 2024-01-24 13:30:08 -08:00
Hui Liu 069cf1d019 use __HIPCC__ for device code 2024-01-24 11:30:01 -08:00
Jiaming Yuan cacb4b1fdd Fix gain calculation in multi-target tree. (#9978) 2024-01-17 13:18:44 +08:00
Hui Liu 1e1e8be3a5 merge latest, Jan 12 2024 2024-01-12 09:57:11 -08:00
Jiaming Yuan c03a4d5088 Check support status for categorical features. (#9946) 2024-01-04 16:51:33 +08:00
Jiaming YuanandPhilip Hyunsu Cho 621348abb3 Fix multi-output with alternating strategies. (#9933)
---------

Co-authored-by: Philip Hyunsu Cho <chohyu01@cs.washington.edu>
2024-01-04 16:41:13 +08:00
Jiaming Yuan a7226c0222 Fix feature names with special characters. (#9923) 2023-12-28 22:45:13 +08:00
Hui Liu 2d7ffbdf3d merge latest changes 2023-12-13 21:06:28 -08:00
Jiaming Yuan fedd9674c8 Implement column sampler in CUDA. (#9785)
- CUDA implementation.
- Extract the broadcasting logic, we will need the context parameter after revamping the collective implementation.
- Some changes to the event loop for fixing a deadlock in CI.
- Move argsort into algorithms.cuh, add support for cuda stream.
2023-11-17 04:29:08 +08:00
Jiaming Yuan 06bdc15e9b [coll] Pass context to various functions. (#9772)
* [coll] Pass context to various functions.

In the future, the `Context` object would be required for collective operations, this PR
passes the context object to some required functions to prepare for swapping out the
implementation.
2023-11-08 09:54:05 +08:00
Hui Liu 8fab17ae8f rm hip.h files 2023-10-30 21:20:28 -07:00
Hui Liu d7f1235b7d Merge branch 'master' into sync-condition-2023Oct11 2023-10-30 13:19:33 -07:00
Jiaming Yuan 6755179e77 [coll] Add nccl. (#9726) 2023-10-28 16:33:58 +08:00
Hui Liu 6762230d9a namespace to reduce code 2023-10-27 10:51:32 -07:00
Hui Liu 4a4b528d54 add namespace aliases to reduce code 2023-10-27 09:11:55 -07:00
Hui Liu 3752b06550 Merge branch 'master' into sync-condition-2023Oct11 2023-10-24 10:46:38 -07:00
Jiaming Yuan 7a02facc9d Serialize expand entry for allgather. (#9702) 2023-10-24 14:33:28 +08:00
Hui Liu 79319dfd4d format 2023-10-23 22:29:48 -07:00
Hui Liu 15421e40d9 enable ROCm on latest XGBoost 2023-10-23 11:07:08 -07:00
Dmitry Razdoburdin ea9f09716b Reorder if-else statements to allow using of cpu branches for sycl-devices (#9682) 2023-10-18 10:55:33 +08:00
Your Name ffbbc9c968 add cuda to hip wrapper 2023-10-17 12:42:37 -07:00
Your Name ea19555474 temp merge, disable 1 line, SetValid 2023-10-12 16:16:44 -07:00
Rong Ou e164d51c43 Improve allgather functions (#9649) 2023-10-12 23:31:43 +08:00
Jiaming Yuan 8c676c889d Remove internal use of gpu_id. (#9568) 2023-09-20 23:29:51 +08:00
Rong Ou c928dd4ff5 Support vertical federated learning with gpu_hist (#9539) 2023-09-03 11:37:11 +08:00
Rong Ou 9bab06cbca Support column split in gpu hist updater (#9384) 2023-08-31 18:09:35 +08:00
Jiaming Yuan ddf2e68821 Use the new DeviceOrd in the linalg module. (#9527) 2023-08-29 13:37:29 +08:00
Jiaming Yuan 942b957eef Fix GPU categorical split memory allocation. (#9529) 2023-08-29 10:06:03 +08:00
Jiaming Yuan 972730cde0 Use matrix for gradient. (#9508)
- Use the `linalg::Matrix` for storing gradients.
- New API for the custom objective.
- Custom objective for multi-class/multi-target is now required to return the correct shape.
- Custom objective for Python can accept arrays with any strides. (row-major, column-major)
2023-08-24 05:29:52 +08:00
Rong Ou 6103dca0bb Support column split in GPU evaluate splits (#9511) 2023-08-23 16:33:43 +08:00
Jiaming Yuan 05d7000096 Handle special characters in JSON model dump. (#9474) 2023-08-14 15:49:00 +08:00
Jiaming Yuan 1caa93221a Use realloc for histogram cache and expose the cache limit. (#9455) 2023-08-10 14:05:27 +08:00
Jiaming Yuan 54029a59af Bound the size of the histogram cache. (#9440)
- A new histogram collection with a limit in size.
- Unify histogram building logic between hist, multi-hist, and approx.
2023-08-08 03:21:26 +08:00
Jiaming Yuan 1332ff787f Unify the code path between local and distributed training. (#9433)
This removes the need for a local histogram space during distributed training, which cuts the cache size by half.
2023-08-03 21:46:36 +08:00