xgboost

Author	SHA1	Message	Date
Tong He	7b74b1b64d	fix additional files note (#4699 ) * fix additional files note * Trigger CI * Trigger CI	2019-07-25 00:37:38 -07:00
Jiaming Yuan	59bc1ef330	Remove VC-2013 support. (#4701 ) * Removing it as it is not fully c++11 compliance.	2019-07-25 01:28:51 -04:00
Philip Hyunsu Cho	2758c5acea	[CI] Fix broken installation of Pandas (#4704 )	2019-07-24 19:03:35 -07:00
Nan Zhu	d5c386ae24	Update CONTRIBUTORS.md	2019-07-24 09:38:47 -07:00
Jiaming Yuan	001aaaee5f	Removed deprecated gpu objectives. (#4690 )	2019-07-20 23:18:34 -04:00
Philip Hyunsu Cho	d4e0a30582	Upgrade dmlc-core submodule (#4688 )	2019-07-20 11:30:05 -07:00
Jiaming Yuan	f0064c07ab	Refactor configuration [Part II]. (#4577 ) * Refactor configuration [Part II]. * General changes: Remove `Init` methods to avoid ambiguity. Remove `Configure(std::map<>)` to avoid redundant copying and prepare for parameter validation. (`std::vector` is returned from `InitAllowUnknown`). ** Add name to tree updaters for easier debugging. * Learner changes: Make `LearnerImpl` the only source of configuration. All configurations are stored and carried out by `LearnerImpl::Configure()`. Remove booster in C API. Originally kept for "compatibility reason", but did not state why. So here we just remove it. Add a `metric_names_` field in `LearnerImpl`. Remove `LazyInit`. Configuration will always be lazy. ** Run `Configure` before every iteration. * Predictor changes: Allocate both cpu and gpu predictor. Remove cpu_predictor from gpu_predictor. `GBTree` is now used to dispatch the predictor. ** Remove some GPU Predictor tests. * IO No IO changes. The binary model format stability is tested by comparing hashing value of save models between two commits	2019-07-20 08:34:56 -04:00
Jiaming Yuan	ad1192e8a3	Remove `silent` in doc. [skip ci] (#4689 )	2019-07-20 05:53:42 -04:00
Nathan Moore	b45258ce66	minor updates to links and grammar (#4673 ) updated links to caret data splitting, xgb.dump(with_stats), and some grammar	2019-07-18 16:56:40 -07:00
Philip Hyunsu Cho	4ef6d216b9	Upgrade dmlc-core submodule (#4674 )	2019-07-17 18:54:49 -07:00
Tong He	8ac8fbef29	[R] Fix CRAN error for Mac OS X (#4672 ) * fix cran error for mac os x * ignore float on windows check for now	2019-07-17 17:55:52 -07:00
Nan Zhu	1595e3f57b	upgrade version num (#4670 ) * upgrade version num * missign changes * fix version script * change versions * rm files * Update CMakeLists.txt	2019-07-17 15:25:35 -07:00
Nan Zhu	01b0c9047c	[jvm-packages] allowing chaining prediction (#4667 ) * add test for chaining prediction * update rabit * Update XGBoostGeneralSuite.scala	2019-07-17 08:50:27 -07:00
koertkuipers	3c506b076e	[jvm-packages] upgrade to Scala 2.12 (#4574 ) * bump scala to 2.12 which requires java 8 and also newer flink and akka * put scala version in artifactId * fix appveyor * fix for scaladoc issue that looks like https://github.com/scala/bug/issues/10509 * fix ci_build * update versions in generate_pom.py * fix generate_pom.py * apache does not have a download for spark 2.4.3 distro using scala 2.12 yet, so for now i use a tgz i put on s3 * Upload spark-2.4.3-bin-scala2.12-hadoop2.7.tgz to our own S3 * Update Dockerfile.jvm_cross * Update Dockerfile.jvm_cross	2019-07-16 08:43:34 -07:00
Oleksandr Pryimak	5544a730f1	Add optional dependencies to setup.py (#4655 )	2019-07-16 17:12:43 +08:00
Mathew Wicks	6323ef94ad	[jvm-packages] update local dev build process (#4640 )	2019-07-15 21:23:06 -07:00
Philip Hyunsu Cho	9975c533c7	Re-organize contributor's guide (#4659 ) * Reorganize contributor's doc * Address comments from @trivialfis * Address @sriramch's comment: include ABI compatibility guarantee * Address @rongou's comment * Postpone ABI compatibility guarantee for now	2019-07-15 20:56:05 -07:00
Oleksandr Pryimak	2973416f2e	[jvm-packages] Fix maven warnings (#4664 ) * exec plugin was missing a version * reportPlugins has been deprecated: see https://maven.apache.org/plugins/maven-site-plugin/maven-3.html#Classic_configuration_Maven_2__3	2019-07-15 20:25:43 -07:00
Matvey Turkov	61f764946f	fixed year to 2019 in conf.py, helpers.h and LICENSE (#4661 )	2019-07-15 12:29:12 -04:00
Mingjie Tang	beb7b295a8	Add tutorial for distributed training and batch prediction with Kubernetes (#4621 ) * provide the readme * update for format * reformat * reformat -2 * update again * update format * update w.r.t yinlou's comments * Add kubernetes tutorial to Table of Contents * Style edit	2019-07-14 23:27:27 -07:00
Nan Zhu	3e339d9557	contribute to community doc (#4646 ) * add community doc * update * update	2019-07-14 21:29:57 -07:00
sriramch	7a388cbf8b	Modify caching allocator/vector and fix issues relating to inability to train large datasets (#4615 )	2019-07-09 18:33:27 +12:00
Xu Xiao	cd1526d3b1	fix auc error in distributed mode caused by unbalanced dataset (#4645 )	2019-07-08 16:01:52 +08:00
Rong Ou	30204b50fe	fix spark tests on machines with many cores (#4634 )	2019-07-07 16:02:56 -07:00
Philip Hyunsu Cho	d333918f5e	[jvm-packages] Expose setMissing method in XGBoostClassificationModel / XGBoostRegressionModel (#4643 )	2019-07-07 16:02:44 -07:00
Philip Hyunsu Cho	1aaf4a679d	Fix early stopping in the Python package (#4638 ) * Fix #4630, #4421: Preserve correct ordering between metrics, and always use last metric for early stopping * Clarify semantics of early stopping in presence of multiple valid sets and metrics * Add a test * Fix lint	2019-07-07 01:01:03 -07:00
Marcos	562d9ae963	Eliminate FutureWarning: Series.base is deprecated (#4337 ) * Remove all references to data.base Should eliminate the deprecation warning in issue #4300 * Fix lint	2019-07-04 21:06:23 -07:00
Jiaming Yuan	d9a47794a5	Fix CPU hist init for sparse dataset. (#4625 ) * Fix CPU hist init for sparse dataset. * Implement sparse histogram cut. * Allow empty features. * Fix windows build, don't use sparse in distributed environment. * Comments. * Smaller threshold. * Fix windows omp. * Fix msvc lambda capture. * Fix MSVC macro. * Fix MSVC initialization list. * Fix MSVC initialization list x2. * Preserve categorical feature behavior. * Rename matrix to sparse cuts. * Reuse UseGroup. * Check for categorical data when adding cut. Co-Authored-By: Philip Hyunsu Cho <chohyu01@cs.washington.edu> * Sanity check. * Fix comments. * Fix comment.	2019-07-04 16:27:03 -07:00
Philip Hyunsu Cho	b7a1f22d24	Empty evaluation list in early stopping should produce meaningful error message (#4633 ) * Empty evaluation list should not break early stopping * Fix lint * Update callback.py	2019-07-04 13:27:18 -07:00
Philip Hyunsu Cho	4df246191f	Add warning when save_model() is called from scikit-learn interface (#4632 )	2019-07-03 23:37:53 -07:00
Philip Hyunsu Cho	96bf91725b	Support ndcg- and map- (#4635 )	2019-07-03 22:51:48 -07:00
Philip Hyunsu Cho	4e9fad74eb	[R] Use built-in label when xgb.DMatrix is given to xgb.cv() (#4631 ) * Use built-in label when xgb.DMatrix is given to xgb.cv() * Add a test * Fix test * Bump version number	2019-07-03 01:32:40 -07:00
Oleksandr Pryimak	986fee6022	pytest tests/python fails if no pandas installed (#4620 ) * _maybe_pandas_xxx should return their arguments unchanged if no pandas installed * Tests should not assume pandas is installed * Mark tests which require pandas as such	2019-07-01 02:54:08 +08:00
Jiaming Yuan	45876bf41b	Fix external memory for get column batches. (#4622 ) * Fix external memory for get column batches. This fixes two bugs: * Use PushCSC for get column batches. * Don't remove the created temporary directory before finishing test. * Check all pages.	2019-06-30 09:56:49 +08:00
Philip Hyunsu Cho	a30176907f	Support Dask 2.0 (#4617 )	2019-06-27 20:42:35 -07:00
Oleksandr Pryimak	923e6c86ba	Add to documentation how to run tests locally (#4610 ) * Add to documentation how to build native unit tests * Add instructions to run Python tests and to use Docker container [skip ci] * Fix link to pytest chapter * Add link to Google Test [skip ci] * Set PYTHONPATH [skip ci] * Revise test_python.sh for running tests locally * Update test_python.sh * Place Docker recommendation notice in a prominent place [skip ci]	2019-06-27 19:02:04 -07:00
Rong Ou	63ec95623d	fix gpu predictor when dmatrix is mismatched with model (#4613 )	2019-06-28 11:03:02 +12:00
Egor Smirnov	4d6590be3c	Optimize ‘hist’ for multi-core CPU (#4529 ) * Initial performance optimizations for xgboost * remove includes * revert float->double * fix for CI * fix for CI * fix for CI * fix for CI * fix for CI * fix for CI * fix for CI * fix for CI * fix for CI * fix for CI * Check existence of _mm_prefetch and __builtin_prefetch * Fix lint * optimizations for CPU * appling comments in review * add some comments, code refactoring * fixing issues in CI * adding runtime checks * remove 1 extra check * remove extra checks in BuildHist * remove checks * add debug info * added debug info * revert changes * added comments * Apply suggestions from code review Co-Authored-By: Philip Hyunsu Cho <chohyu01@cs.washington.edu> * apply review comments * Remove unused function CreateNewNodes() * Add descriptive comment on node_idx variable in QuantileHistMaker::Builder::BuildHistsBatch()	2019-06-27 11:33:49 -07:00
Nan Zhu	abffbe014e	[jvm-packages] delete all constraints from spark layer about obj and eval metrics and handle error in jvm layer (#4560 ) * temp * prediction part * remove supported* * add for test * fix param name * add rabit * update rabit * return value of rabit init * eliminate compilation warnings * update rabit * shutdown * update rabit again * check sparkcontext shutdown * fix logic * sleep * fix tests * test with relaxed threshold * create new thread each time * stop for job quitting * udpate rabit * update rabit * update rabit * update git modules	2019-06-27 08:47:37 -07:00
Philip Hyunsu Cho	dd01f7c4f5	Use Sphinx 2.1+ to compile documentation [skip ci] (#4609 )	2019-06-26 16:26:22 -07:00
Philip Hyunsu Cho	cd3a3f99da	Fix doc for customized objective/metric [skip ci] (#4608 )	2019-06-26 13:40:34 -07:00
Jiaming Yuan	5b2f805e74	Doc and demo for customized metric and obj. (#4598 ) Co-Authored-By: Theodore Vasiloudis <theodoros.vasiloudis@gmail.com>	2019-06-26 16:13:12 +08:00
Jiaming Yuan	8bdf15120a	Implement tree model dump with code generator. (#4602 ) * Implement tree model dump with a code generator. * Split up generators. * Implement graphviz generator. * Use pattern matching. * [Breaking] Return a Source in `to_graphviz` instead of Digraph in Python package. Co-Authored-By: Philip Hyunsu Cho <chohyu01@cs.washington.edu>	2019-06-26 15:20:44 +08:00
Nan Zhu	fe2de6f415	[jvm-packages]fix silly bug in feature scoring (#4604 )	2019-06-25 20:49:01 -07:00
Philip Hyunsu Cho	1f98f18cb8	Add instruction to run formatting checks locally [skip ci] (#4591 )	2019-06-24 00:09:09 -07:00
Jiaming Yuan	2cff735126	Update doc for feature constraints and `n_gpus`. (#4596 ) * Update doc for feature constraints. * Fix some warnings. * Clean up doc for `n_gpus`.	2019-06-23 14:37:22 +08:00
Andy Adinets	9fa29ad753	Set reg_lambda=1e-5 for scikit-learn-like random forest classes. (#4558 )	2019-06-22 08:02:13 +12:00
Philip Hyunsu Cho	30e1cb4e9e	Fix docstring for XGBModel.predict() [skip ci] (#4592 )	2019-06-21 12:44:42 +08:00
Rong Ou	77fc28427d	fix benchmark_tree.py (#4593 )	2019-06-21 11:51:48 +12:00
Jiaming Yuan	9494950ee7	Address some sphinx warnings and errors, add doc for building doc. (#4589 )	2019-06-20 15:07:36 -07:00

1 2 3 4 5 ...

3926 Commits