Practical Machine Learning in R

Rs. 9,555
  • Authors: Fred Nwanganga, Mike Chapple
  • ISBN: 9781119591511
  • Publisher: Wiley Publishing
  • Publication Date: July 06, 2020
  • Format: Paperback – 464 pages
  • Language: English

Out of stock



Share
Description

Guides professionals and students through the rapidly growing field of machine learning
with hands-on examples in the popular R programming language

Machine Learning—a branch of Artificial Intelligence (AI) which enables computers to improve their results and learn new approaches without explicit instructions—allows organizations to reveal patterns in their data and incorporate predictive analytics into their decision-making process. Practical Machine Learning in R provides a hands-on approach to solving business problems with intelligent, self-learning computer algorithms.

Bestselling author and data analytics experts Fred Nwanganga and Mike Chapple explain what machine learning is, demonstrate its organizational benefits, and provide hands-on examples created in the R programming language. A perfect guide for professional self-taught learners or students in an introductory machine learning course, this reader-friendly book illustrates the numerous real-world business uses of machine learning approaches. Clear and detailed chapters cover data wrangling, R programming with the popular RStudio tool, classification and regression techniques, performance evaluation, and more.

  • Explores data management techniques, including data collection, exploration and dimensionality reduction
  • Covers unsupervised learning, where readers identify and summarize patterns using approaches such as apriori, eclat and clustering
  • Describes the principles behind the Nearest Neighbor, Decision Tree and Naive Bayes classification techniques
  • Explains how to evaluate and choose the right model, as well as how to improve model performance using ensemble methods such as Random Forest and XGBoost

Practical Machine Learning in R is a must-have guide for business analysts, data scientists, and other professionals interested in leveraging the power of AI to solve business problems, as well as students and independent learners seeking to enter the field.

Table of contents
    1.  About the Authors vii
    2. About the Technical Editors ix
    3. Acknowledgments xi
    4. Introduction xxi
  1. Part I: Getting Started 1
    1. Chapter 1 What is Machine Learning? 3
      1. Discovering Knowledge in Data 5
      2. Introducing Algorithms 5
      3. Artificial Intelligence, Machine Learning, and Deep Learning 6
      4. Machine Learning Techniques 7
      5. Supervised Learning 8
      6. Unsupervised Learning 12
      7. Model Selection 14
      8. Classification Techniques 14
      9. Regression Techniques 15
      10. Similarity Learning Techniques 16
      11. Model Evaluation 16
      12. Classification Errors 17
      13. Regression Errors 19
      14. Types of Error 20
      15. Partitioning Datasets 22
      16. Holdout Method 23
      17. Cross-Validation Methods 23
      18. Exercises 24
    2. Chapter 2 Introduction to R and RStudio 25
      1. Welcome to R 26
      2. R and RStudio Components 27
      3. The R Language 27
      4. RStudio 28
      5. RStudio Desktop 28
      6. RStudio Server 29
      7. Exploring the RStudio
      8. Environment 29
      9. R Packages 38
      10. The CRAN Repository 38
      11. Installing Packages 38
      12. Loading Packages 39
      13. Package Documentation 40
      14. Writing and Running an R Script 41
      15. Data Types in R 44
      16. Vectors 45
      17. Testing Data Types 47
      18. Converting Data Types 50
      19. Missing Values 51
      20. Exercises 52
    3. Chapter 3 Managing Data 53
      1. The Tidyverse 54
      2. Data Collection 55
      3. Key Considerations 55
      4. Collecting Ground Truth Data 55
      5. Data Relevance 55
      6. Quantity of Data 56
      7. Ethics 56
      8. Importing the Data 56
      9. Reading Comma-Delimited Files 56
      10. Reading Other Delimited Files 60
      11. Data Exploration 60
      12. Describing the Data 61
      13. Instance 61
      14. Feature 61
      15. Dimensionality 62
      16. Sparsity and Density 62
      17. Resolution 62
      18. Descriptive Statistics 63
      19. Visualizing the Data 69
      20. Comparison 69
      21. Relationship 70
      22. Distribution 72
      23. Composition 73
      24. Data Preparation 74
      25. Cleaning the Data 75
      26. Missing Values 75
      27. Noise 79
      28. Outliers 81
      29. Class Imbalance 82
      30. Transforming the Data 84
      31. Normalization 84
      32. Discretization 89
      33. Dummy Coding 89
      34. Reducing the Data 92
      35. Sampling 92
      36. Dimensionality Reduction 99
      37. Exercises 100
  2. Part II: Regression 101
    1. Chapter 4 Linear Regression 103
      1. Bicycle Rentals and Regression 104
      2. Relationships Between Variables 106
      3. Correlation 106
      4. Regression 114
      5. Simple Linear Regression 115
      6. Ordinary Least Squares Method 116
      7. Simple Linear Regression Model 119
      8. Evaluating the Model 120
      9. Residuals 121
      10. Coefficients 121
      11. Diagnostics 122
      12. Multiple Linear Regression 124
      13. The Multiple Linear Regression Model 124
      14. Evaluating the Model 125
      15. Residual Diagnostics 127
      16. Influential Point Analysis 130
      17. Multicollinearity 133
      18. Improving the Model 135
      19. Considering Nonlinear Relationships 135
      20. Considering Categorical Variables 137
      21. Considering Interactions Between Variables 139
      22. Selecting the Important Variables 141
      23. Strengths and Weaknesses 146
      24. Case Study: Predicting Blood Pressure 147
      25. Importing the Data 148
      26. Exploring the Data 149
      27. Fitting the Simple Linear Regression Model 151
      28. Fitting the Multiple Linear Regression Model 152
      29. Exercises 161
    2. Chapter 5 Logistic Regression 165
      1. Prospecting for Potential Donors 166
      2. Classifi cation 169
      3. Logistic Regression 170
      4. Odds Ratio 172
      5. Binomial Logistic Regression Model 176
      6. Dealing with Missing Data 178
      7. Dealing with Outliers 182
      8. Splitting the Data 187
      9. Dealing with Class Imbalance 188
      10. Training a Model 190
      11. Evaluating the Model 190
      12. Coeffi cients 193
      13. Diagnostics 195
      14. Predictive Accuracy 195
      15. Improving the Model 198
      16. Dealing with Multicollinearity 198
      17. Choosing a Cutoff Value 205
      18. Strengths and Weaknesses 206
      19. Case Study: Income Prediction 207
      20. Importing the Data 208
      21. Exploring and Preparing the Data 208
      22. Training the Model 212
      23. Evaluating the Model 215
      24. Exercises 216
  3. Part III: Classification 221
    1. Chapter 6 k-Nearest Neighbors 223
      1. Detecting Heart Disease 224
      2. k-Nearest Neighbors 226
      3. Finding the Nearest Neighbors 228
      4. Labeling Unlabeled Data 230
      5. Choosing an Appropriate k 231
      6. k-Nearest Neighbors Model 232
      7. Dealing with Missing Data 234
      8. Normalizing the Data 234
      9. Dealing with Categorical Features 235
      10. Splitting the Data 237
      11. Classifying Unlabeled Data 237
      12. Evaluating the Model 238
      13. Improving the Model 239
      14. Strengths and Weaknesses 241
      15. Case Study: Revisiting the Donor Dataset 241
      16. Importing the Data 241
      17. Exploring and Preparing the Data 242
      18. Dealing with Missing Data 243
      19. Normalizing the Data 245
      20. Splitting and Balancing the Data 246
      21. Building the Model 248
      22. Evaluating the Model 248
      23. Exercises 249
    2. Chapter 7 Naïve Bayes 251
      1. Classifying Spam Email 252
      2. Naïve Bayes 253
      3. Probability 254
      4. Joint Probability 255
      5. Conditional Probability 256
      6. Classification with Naïve Bayes 257
      7. Additive Smoothing 261
      8. Naïve Bayes Model 263
      9. Splitting the Data 266
      10. Training a Model 267
      11. Evaluating the Model 267
      12. Strengths and Weaknesses of the Naïve Bayes Classifier 269
      13. Case Study: Revisiting the Heart Disease Detection Problem 269
      14. Importing the Data 270
      15. Exploring and Preparing the Data 270
      16. Building the Model 272
      17. Evaluating the Model 273
      18. Exercises 274
    3. Chapter 8 Decision Trees 277
      1. Predicting Build Permit Decisions 278
      2. Decision Trees 279
      3. Recursive Partitioning 281
      4. Entropy 285
      5. Information Gain 286
      6. Gini Impurity 290
      7. Pruning 290
      8. Building a Classification Tree Model 291
      9. Splitting the Data 294
      10. Training a Model 295
      11. Evaluating the Model 295
      12. Strengths and Weaknesses of the Decision Tree Model 298
      13. Case Study: Revisiting the Income Prediction Problem 299
      14. Importing the Data 300
      15. Exploring and Preparing the Data 300
      16. Building the Model 302
      17. Evaluating the Model 302
      18. Exercises 304
  4. Part IV: Evaluating and Improving Performance 305
    1. Chapter 9 Evaluating Performance 307
      1. Estimating Future Performance 308
      2. Cross-Validation 311
      3. k-Fold Cross-Validation 311
      4. Leave-One-Out Cross-Validation 315
      5. Random Cross-Validation 316
      6. Bootstrap Sampling 318
      7. Beyond Predictive Accuracy 321
      8. Kappa 323
      9. Precision and Recall 326
      10. Sensitivity and Specificity 328
      11. Visualizing Model Performance 332
      12. Receiver Operating Characteristic Curve 333
      13. Area Under the Curve 336
      14. Exercises 339
    2. Chapter 10 Improving Performance 341
      1. Parameter Tuning 342
      2. Automated Parameter Tuning 342
      3. Customized Parameter Tuning 348
      4. Ensemble Methods 354
      5. Bagging 355
      6. Boosting 358
      7. Stacking 361
      8. Exercises 366
  5. Part V: Unsupervised Learning 367
    1. Chapter 11 Discovering Patterns with Association Rules 369
      1. Market Basket Analysis 370
      2. Association Rules 371
      3. Identifying Strong Rules 373
      4. Support 373
      5. Confi dence 373
      6. Lift 374
      7. The Apriori Algorithm 374
      8. Discovering Association Rules 376
      9. Generating the Rules 377
      10. Evaluating the Rules 382
      11. Strengths and Weaknesses 386
      12. Case Study: Identifying Grocery Purchase Patterns 386
      13. Importing the Data 387
      14. Exploring and Preparing the Data 387
      15. Generating the Rules 389
      16. Evaluating the Rules 389
      17. Exercises 392
      18. Notes 393
    2. Chapter 12 Grouping Data with Clustering 395
      1. Clustering 396
      2. k-Means Clustering 399
      3. Segmenting Colleges with k-Means Clustering 403
      4. Creating the Clusters 404
      5. Analyzing the Clusters 407
      6. Choosing the Right Number of Clusters 409
      7. The Elbow Method 409
      8. The Average Silhouette Method 411
      9. The Gap Statistic 412
      10. Strengths and Weaknesses of k-Means Clustering 414
      11. Case Study: Segmenting Shopping Mall Customers 415
      12. Exploring and Preparing the Data 415
      13. Clustering the Data 416
      14. Evaluating the Clusters 418
      15. Exercises 420
      16. Notes 420
    3. Index 421
Authors Biography

FRED NWANGANGA, PHD, is an assistant teaching professor of business analytics at the University of Notre Dame’s Mendoza College of Business. He has over 15 years of technology leadership experience.

MIKE CHAPPLE, PHD, is associate teaching professor of information technology, analytics, and operations at the Mendoza College of Business. Mike is a bestselling author of over 25 books, and he currently serves as academic director of the University’s Master of Science in Business Analytics program.

Additional information
Weight0.920 kg
Reviews

There are no reviews yet.

Only logged in customers who have purchased this product may leave a review.