Go to file

Avinash Barnwal dcf439932a

Add Accelerated Failure Time loss for survival analysis task (#4763 )

* [WIP] Add lower and upper bounds on the label for survival analysis

* Update test MetaInfo.SaveLoadBinary to account for extra two fields

* Don't clear qids_ for version 2 of MetaInfo

* Add SetInfo() and GetInfo() method for lower and upper bounds

* changes to aft

* Add parameter class for AFT; use enum's to represent distribution and event type

* Add AFT metric

* changes to neg grad to grad

* changes to binomial loss

* changes to overflow

* changes to eps

* changes to code refactoring

* changes to code refactoring

* changes to code refactoring

* Re-factor survival analysis

* Remove aft namespace

* Move function bodies out of AFTNormal and AFTLogistic, to reduce clutter

* Move function bodies out of AFTLoss, to reduce clutter

* Use smart pointer to store AFTDistribution and AFTLoss

* Rename AFTNoiseDistribution enum to AFTDistributionType for clarity

The enum class was not a distribution itself but a distribution type

* Add AFTDistribution::Create() method for convenience

* changes to extreme distribution

* changes to extreme distribution

* changes to extreme

* changes to extreme distribution

* changes to left censored

* deleted cout

* changes to x,mu and sd and code refactoring

* changes to print

* changes to hessian formula in censored and uncensored

* changes to variable names and pow

* changes to Logistic Pdf

* changes to parameter

* Expose lower and upper bound labels to R package

* Use example weights; normalize log likelihood metric

* changes to CHECK

* changes to logistic hessian to standard formula

* changes to logistic formula

* Comply with coding style guideline

* Revert back Rabit submodule

* Revert dmlc-core submodule

* Comply with coding style guideline (clang-tidy)

* Fix an error in AFTLoss::Gradient()

* Add missing files to amalgamation

* Address @RAMitchell's comment: minimize future change in MetaInfo interface

* Fix lint

* Fix compilation error on 32-bit target, when size_t == bst_uint

* Allocate sufficient memory to hold extra label info

* Use OpenMP to speed up

* Fix compilation on Windows

* Address reviewer's feedback

* Add unit tests for probability distributions

* Make Metric subclass of Configurable

* Address reviewer's feedback: Configure() AFT metric

* Add a dummy test for AFT metric configuration

* Complete AFT configuration test; remove debugging print

* Rename AFT parameters

* Clarify test comment

* Add a dummy test for AFT loss for uncensored case

* Fix a bug in AFT loss for uncensored labels

* Complete unit test for AFT loss metric

* Simplify unit tests for AFT metric

* Add unit test to verify aggregate output from AFT metric

* Use EXPECT_* instead of ASSERT_*, so that we run all unit tests

* Use aft_loss_param when serializing AFTObj

This is to be consistent with AFT metric

* Add unit tests for AFT Objective

* Fix OpenMP bug; clarify semantics for shared variables used in OpenMP loops

* Add comments

* Remove AFT prefix from probability distribution; put probability distribution in separate source file

* Add comments

* Define kPI and kEulerMascheroni in probability_distribution.h

* Add probability_distribution.cc to amalgamation

* Remove unnecessary diff

* Address reviewer's feedback: define variables where they're used

* Eliminate all INFs and NANs from AFT loss and gradient

* Add demo

* Add tutorial

* Fix lint

* Use 'survival:aft' to be consistent with 'survival:cox'

* Move sample data to demo/data

* Add visual demo with 1D toy data

* Add Python tests

Co-authored-by: Philip Cho <chohyu01@cs.washington.edu>

2020-03-25 13:52:51 -07:00

.github

Display Sponsor button, link to OpenCollective (#5325 )

2020-02-19 01:58:21 -08:00

amalgamation

Add Accelerated Failure Time loss for survival analysis task (#4763 )

2020-03-25 13:52:51 -07:00

cmake

[R-package] changed FindLibR to take advantage of CMake cache (#5427 )

2020-03-20 03:32:15 +08:00

cub @ b20808b1b0

Update cub submodule again (fixes GPU build) (#2599 )

2017-08-13 22:14:40 +12:00

demo

Add Accelerated Failure Time loss for survival analysis task (#4763 )

2020-03-25 13:52:51 -07:00

dev

Add release note for 1.0.0 in NEWS.md (#5329 )

2020-03-03 21:35:43 -08:00

dmlc-core @ 552f7de748

Add CMake option to run Undefined Behavior Sanitizer (UBSan) (#5211 )

2020-01-20 16:57:44 +08:00

doc

Add Accelerated Failure Time loss for survival analysis task (#4763 )

2020-03-25 13:52:51 -07:00

include/xgboost

Add Accelerated Failure Time loss for survival analysis task (#4763 )

2020-03-25 13:52:51 -07:00

jvm-packages

Add number of columns to native data iterator. (#5202 )

2020-02-25 23:42:01 +08:00

make

Use DART tree weights when computing SHAPs (#5050 )

2019-12-03 19:55:53 +08:00

plugin

Add link to GPU documentation (#5437 )

2020-03-24 09:29:29 +13:00

python-package

[dask] Fix missing value for scikit-learn interface. (#5435 )

2020-03-20 10:56:01 -04:00

R-package

Add Accelerated Failure Time loss for survival analysis task (#4763 )

2020-03-25 13:52:51 -07:00

rabit @ 2f7fcff4d7

Update Rabit (#5237 )

2020-01-28 02:05:01 -08:00

src

Add Accelerated Failure Time loss for survival analysis task (#4763 )

2020-03-25 13:52:51 -07:00

tests

Add Accelerated Failure Time loss for survival analysis task (#4763 )

2020-03-25 13:52:51 -07:00

.clang-tidy

Don't use modernize-use-trailing-return-type. (#5169 )

2019-12-29 20:18:23 +08:00

.editorconfig

Added configuration for python into .editorconfig (#3494 )

2018-07-23 00:24:10 -07:00

.gitignore

Ignore gdb_history. [skip ci] (#5257 )

2020-02-02 20:40:09 +08:00

.gitmodules

Upgrading to NCCL2 (#3404 )

2018-07-10 00:42:15 -07:00

.travis.yml

Make pip install xgboost*.tar.gz work by fixing build-python.sh (#5241 )

2020-01-28 23:18:23 -08:00

appveyor.yml

Remove VC-2013 support. (#4701 )

2019-07-25 01:28:51 -04:00

CITATION

simplify software citation (#2912 )

2017-12-01 02:58:13 -08:00

CMakeLists.txt

Adding static library option (#5397 )

2020-03-10 18:22:15 +08:00

CONTRIBUTORS.md

Update affiliation of @hcho3 (#5292 )

2020-02-06 20:58:39 -08:00

Jenkinsfile

Revert "Enable rabit test (#5358 )" (#5377 )

2020-02-29 04:25:03 +08:00

Jenkinsfile-win64

[CI] Upload master branch artifacts to S3 root [skip ci] (#4979 )

2019-10-23 22:39:04 -07:00

LICENSE

fixed year to 2019 in conf.py, helpers.h and LICENSE (#4661 )

2019-07-15 12:29:12 -04:00

Makefile

Rewrite setup.py. (#5271 )

2020-02-04 13:35:42 +08:00

NEWS.md

Add release note for 1.0.0 in NEWS.md (#5329 )

2020-03-03 21:35:43 -08:00

README.md

Update README.md (#5346 )

2020-02-23 02:52:37 +08:00

README.md

eXtreme Gradient Boosting

Community | Documentation | Resources | Contributors | Release Notes

XGBoost is an optimized distributed gradient boosting library designed to be highly efficient, flexible and portable. It implements machine learning algorithms under the Gradient Boosting framework. XGBoost provides a parallel tree boosting (also known as GBDT, GBM) that solve many data science problems in a fast and accurate way. The same code runs on major distributed environment (Kubernetes, Hadoop, SGE, MPI, Dask) and can solve problems beyond billions of examples.

License

Contribute to XGBoost

XGBoost has been developed and used by a group of active community members. Your help is very valuable to make the package better for everyone. Checkout the Community Page.

Reference

Tianqi Chen and Carlos Guestrin. XGBoost: A Scalable Tree Boosting System. In 22nd SIGKDD Conference on Knowledge Discovery and Data Mining, 2016
XGBoost originates from research project at University of Washington.