xgboost

Author	SHA1	Message	Date
Jiaming Yuan	ccd30e4491	Fix non-openmp build. (#5566 ) * Add test to Jenkins. * Fix threading utils tests. * Require thread library.	2020-04-20 12:16:38 +08:00
Rory Mitchell	d6d1035950	gpu_hist performance fixes (#5558 ) * Remove unnecessary cuda API calls * Fix histogram memory growth	2020-04-19 12:21:13 +12:00
Jiaming Yuan	e1f22baf8c	Fix slice and get info. (#5552 )	2020-04-18 18:00:13 +08:00
Jiaming Yuan	c245eb8755	Fix r interaction constraints (#5543 ) * Unify the parsing code. * Cleanup.	2020-04-18 06:53:51 +08:00
Jiaming Yuan	93df871c8c	Assert matching length of evaluation inputs. (#5540 )	2020-04-18 06:52:55 +08:00
Jiaming Yuan	c69a19e2b1	Fix skl nan tag. (#5538 )	2020-04-18 06:52:17 +08:00
ShvetsKS	a2d86b8e4b	Optimizations for RNG in InitData kernel (#5522 ) * optimizations for subsampling in InitData * optimizations for subsampling in InitData Co-authored-by: SHVETS, KIRILL <kirill.shvets@intel.com>	2020-04-16 18:24:32 +03:00
Rory Mitchell	e268fb0093	Use thrust functions instead of custom functions (#5544 )	2020-04-16 21:41:16 +12:00
Jiaming Yuan	468b1594d3	Fix CLI model IO. (#5535 ) * Add test for comparing Python and CLI training result.	2020-04-16 07:48:47 +08:00
Philip Hyunsu Cho	0676a19e70	[jvm-packages] [CI] Publish XGBoost4J JARs with Scala 2.11 and 2.12 (#5539 )	2020-04-15 09:32:02 -07:00
Philip Hyunsu Cho	ec02f40d42	[CI] Use Ubuntu 18.04 LTS in JVM CI, because 19.04 is EOL (#5537 )	2020-04-15 07:32:46 -07:00
Jiaming Yuan	8b04736b81	[dask] dask cudf inplace prediction. (#5512 ) * Add inplace prediction for dask-cudf. * Remove Dockerfile.release, since it's not used anywhere * Use Conda exclusively in CUDF and GPU containers * Improve cupy memory copying. * Add skip marks to tests. * Add mgpu-cudf category on the CI to run all distributed tests. Co-authored-by: Hyunsu Cho <chohyu01@cs.washington.edu>	2020-04-15 18:15:51 +08:00
Rory Mitchell	ca4e05660e	Purge device_helpers.cuh (#5534 ) * Simplifications with caching_device_vector * Purge device helpers	2020-04-15 21:51:56 +12:00
Philip Hyunsu Cho	1b1969f20d	[jvm-packages] [CI] Create a Maven repository to host SNAPSHOT JARs (#5533 )	2020-04-14 19:33:32 -07:00
Philip Hyunsu Cho	88b64c8162	Ensure that configured dmlc/build_config.h is picked up by Rabit and XGBoost (#5514 ) * Ensure that configured header (build_config.h) from dmlc-core is picked up by Rabit and XGBoost * Check which Rabit target is being used * Use CMake 3.13 in all Jenkins tests * Upgrade CMake in Travis CI * Install CMake using Kitware installer * Remove existing CMake (3.12.4)	2020-04-11 23:48:28 -07:00
Liang-Chi Hsieh	449ab79e0c	[CI] Use devtoolset-6 because devtoolset-4 is EOL and no longer available (#5506 ) * Use devtoolset-6. * [CI] Use devtoolset-6 because devtoolset-4 is EOL and no longer available * CUDA 9.0 doesn't work with devtoolset-6; use devtoolset-4 for GPU build only Co-authored-by: Hyunsu Cho <chohyu01@cs.washington.edu>	2020-04-11 19:49:06 -07:00
Jiaming Yuan	a3db79df22	Remove makefiles. (#5513 )	2020-04-11 13:25:53 +08:00
Rory Mitchell	093e2227e3	Serialise booster after training to reset state (#5484 ) * Serialise booster after training to reset state * Prevent process_type being set on load * Check for correct updater sequence	2020-04-11 16:27:12 +12:00
Jiaming Yuan	1334aca437	Fix github merge. (#5509 )	2020-04-10 22:17:38 +08:00
Jiaming Yuan	7d52c0b8c2	Requires setting leaf stat when expanding tree. (#5501 ) * Fix GPU Hist feature importance.	2020-04-10 12:27:03 +08:00
Jiaming Yuan	dc2950fd90	Fix checking booster. (#5505 ) * Use `get_params()` instead of `getattr` intrinsic.	2020-04-10 12:21:21 +08:00
Jiaming Yuan	6671b42dd4	Use ellpack for prediction only when sparsepage doesn't exist. (#5504 )	2020-04-10 12:15:46 +08:00
Jiaming Yuan	0012f2ef93	Upgrade clang-tidy on CI. (#5469 ) * Correct all clang-tidy errors. * Upgrade clang-tidy to 10 on CI. Co-authored-by: Hyunsu Cho <chohyu01@cs.washington.edu>	2020-04-05 04:42:29 +08:00
Philip Hyunsu Cho	5fc5ec539d	Implement robust regularization in 'survival:aft' objective (#5473 ) * Robust regularization of AFT gradient and hessian * Fix AFT doc; expose it to tutorial TOC * Apply robust regularization to uncensored case too * Revise unit test slightly * Fix lint * Update test_survival.py * Use GradientPairPrecise * Remove unused variables	2020-04-04 12:21:24 -07:00
Jiaming Yuan	86beb68ce8	Implement host span. (#5459 )	2020-04-03 10:37:51 +08:00
Jiaming Yuan	459b175dc6	Split up test helpers header. (#5455 )	2020-04-03 10:36:53 +08:00
Jiaming Yuan	c218d8ffbf	Enable parameter validation for skl. (#5477 )	2020-04-03 10:23:58 +08:00
Jiaming Yuan	29c6ad943a	Prevent copying SimpleDMatrix. (#5453 ) * Set default dtor for SimpleDMatrix to initialize default copy ctor, which is deleted due to unique ptr. * Remove commented code. * Remove warning for calling host function (std::max). * Remove warning for initialization order. * Remove warning for unused variables.	2020-04-02 07:01:49 +08:00
Jiaming Yuan	e86030c360	Update dmlc-core. (#5466 ) * Copy dmlc travis script to XGBoost.	2020-04-02 04:16:39 +08:00
Jiaming Yuan	babcb996e7	Reduce span check overhead. (#5464 )	2020-04-01 22:07:24 +08:00
Rory Mitchell	15f40e51e9	Add support for dlpack, expose python docs for DeviceQuantileDMatrix (#5465 )	2020-04-01 23:34:32 +13:00
Jiaming Yuan	6601a641d7	Thread safe, inplace prediction. (#5389 ) Normal prediction with DMatrix is now thread safe with locks. Added inplace prediction is lock free thread safe. When data is on device (cupy, cudf), the returned data is also on device. * Implementation for numpy, csr, cudf and cupy. * Implementation for dask. * Remove sync in simple dmatrix.	2020-03-30 15:35:28 +08:00
ShvetsKS	27a8e36fc3	Reducing memory consumption for 'hist' method on CPU (#5334 )	2020-03-28 14:45:52 +13:00
Rory Mitchell	13b10a6370	Device dmatrix (#5420 )	2020-03-28 14:42:21 +13:00
Jiaming Yuan	780de49ddb	Resolve travis failure. (#5445 ) * Install dependencies by pip.	2020-03-27 19:37:58 +08:00
Jiaming Yuan	4942da64ae	Refactor tests with data generator. (#5439 )	2020-03-27 06:44:44 +08:00
Avinash Barnwal	dcf439932a	Add Accelerated Failure Time loss for survival analysis task (#4763 ) * [WIP] Add lower and upper bounds on the label for survival analysis * Update test MetaInfo.SaveLoadBinary to account for extra two fields * Don't clear qids_ for version 2 of MetaInfo * Add SetInfo() and GetInfo() method for lower and upper bounds * changes to aft * Add parameter class for AFT; use enum's to represent distribution and event type * Add AFT metric * changes to neg grad to grad * changes to binomial loss * changes to overflow * changes to eps * changes to code refactoring * changes to code refactoring * changes to code refactoring * Re-factor survival analysis * Remove aft namespace * Move function bodies out of AFTNormal and AFTLogistic, to reduce clutter * Move function bodies out of AFTLoss, to reduce clutter * Use smart pointer to store AFTDistribution and AFTLoss * Rename AFTNoiseDistribution enum to AFTDistributionType for clarity The enum class was not a distribution itself but a distribution type * Add AFTDistribution::Create() method for convenience * changes to extreme distribution * changes to extreme distribution * changes to extreme * changes to extreme distribution * changes to left censored * deleted cout * changes to x,mu and sd and code refactoring * changes to print * changes to hessian formula in censored and uncensored * changes to variable names and pow * changes to Logistic Pdf * changes to parameter * Expose lower and upper bound labels to R package * Use example weights; normalize log likelihood metric * changes to CHECK * changes to logistic hessian to standard formula * changes to logistic formula * Comply with coding style guideline * Revert back Rabit submodule * Revert dmlc-core submodule * Comply with coding style guideline (clang-tidy) * Fix an error in AFTLoss::Gradient() * Add missing files to amalgamation * Address @RAMitchell's comment: minimize future change in MetaInfo interface * Fix lint * Fix compilation error on 32-bit target, when size_t == bst_uint * Allocate sufficient memory to hold extra label info * Use OpenMP to speed up * Fix compilation on Windows * Address reviewer's feedback * Add unit tests for probability distributions * Make Metric subclass of Configurable * Address reviewer's feedback: Configure() AFT metric * Add a dummy test for AFT metric configuration * Complete AFT configuration test; remove debugging print * Rename AFT parameters * Clarify test comment * Add a dummy test for AFT loss for uncensored case * Fix a bug in AFT loss for uncensored labels * Complete unit test for AFT loss metric * Simplify unit tests for AFT metric * Add unit test to verify aggregate output from AFT metric * Use EXPECT_* instead of ASSERT_, so that we run all unit tests Use aft_loss_param when serializing AFTObj This is to be consistent with AFT metric * Add unit tests for AFT Objective * Fix OpenMP bug; clarify semantics for shared variables used in OpenMP loops * Add comments * Remove AFT prefix from probability distribution; put probability distribution in separate source file * Add comments * Define kPI and kEulerMascheroni in probability_distribution.h * Add probability_distribution.cc to amalgamation * Remove unnecessary diff * Address reviewer's feedback: define variables where they're used * Eliminate all INFs and NANs from AFT loss and gradient * Add demo * Add tutorial * Fix lint * Use 'survival:aft' to be consistent with 'survival:cox' * Move sample data to demo/data * Add visual demo with 1D toy data * Add Python tests Co-authored-by: Philip Cho <chohyu01@cs.washington.edu>	2020-03-25 13:52:51 -07:00
sriramch	d2231fc840	Ranking metric acceleration on the gpu (#5398 )	2020-03-22 19:38:48 +13:00
Jiaming Yuan	cd7d6f7d59	[dask] Fix missing value for scikit-learn interface. (#5435 )	2020-03-20 10:56:01 -04:00
Jiaming Yuan	abca9908ba	Support pandas SparseArray. (#5431 )	2020-03-20 21:40:22 +08:00
Jiaming Yuan	760d5d0c3c	[dask] Accept other inputs for prediction. (#5428 ) * Returns a series when input is dataframe. * Merge assert client.	2020-03-19 17:05:55 +08:00
Jiaming Yuan	b51124c158	[dask] Enable gridsearching with skl. (#5417 )	2020-03-16 04:51:51 +08:00
Jiaming Yuan	761a5dbdfc	[dask] Honor `nthreads` from dask worker. (#5414 )	2020-03-16 04:51:24 +08:00
Jiaming Yuan	21b671aa06	[dask] Order the prediction result. (#5416 )	2020-03-15 19:34:04 +08:00
Jiaming Yuan	ab7a46a1a4	Check whether current updater can modify a tree. (#5406 ) * Check whether current updater can modify a tree. * Fix tree model JSON IO for pruned trees.	2020-03-14 09:24:08 +08:00
Rory Mitchell	b745b7acce	Fix memory usage of device sketching (#5407 )	2020-03-14 13:43:24 +13:00
Rory Mitchell	3ad4333b0e	Partial rewrite EllpackPage (#5352 )	2020-03-11 10:15:53 +13:00
Rory Mitchell	a38e7bd19c	Sketching from adapters (#5365 ) * Sketching from adapters * Add weights test	2020-03-07 21:07:58 +13:00
Jiaming Yuan	0dd97c206b	Move thread local entry into Learner. (#5396 ) * Move thread local entry into Learner. This is an attempt to workaround CUDA context issue in static variable, where the CUDA context can be released before device vector. * Add PredictionEntry to thread local entry. This eliminates one copy of prediction vector. * Don't define CUDA C API in a namespace.	2020-03-07 15:37:39 +08:00
Jiaming Yuan	8d06878bf9	Deterministic GPU histogram. (#5361 ) * Use pre-rounding based method to obtain reproducible floating point summation. * GPU Hist for regression and classification are bit-by-bit reproducible. * Add doc. * Switch to thrust reduce for `node_sum_gradient`.	2020-03-04 15:13:28 +08:00

1 2 3 4 5 ...

525 Commits