Mastering Python Data Analysis

Rs. 11,645
  • Authors: Magnus Vilhelm Persson, Luiz Felipe Martins
  • ISBN: 9781783553297
  • Publisher: Packt Publishing
  • Publication Date: June 30, 2016
  • Format: Paperback – 284 pages
  • Language: English

Out of stock



Share
Description

Become an expert at using Python for advanced statistical analysis of data using real-world examples About This Book

  • Clean, format, and explore data using graphical and numerical summaries
  • Leverage the IPython environment to efficiently analyze data with Python
  • Packed with easy-to-follow examples to develop advanced computational skills for the analysis of complex data Who This Book Is For If you are a competent Python developer who wants to take your data analysis skills to the next level by solving complex problems, then this advanced guide is for you. Familiarity with the basics of applying Python libraries to data sets is assumed. What You Will Learn
  • Read, sort, and map various data into Python and Pandas
  • Recognise patterns so you can understand and explore data
  • Use statistical models to discover patterns in data
  • Review classical statistical inference using Python, Pandas, and SciPy
  • Detect similarities and differences in data with clustering
  • Clean your data to make it useful
  • Work in Jupyter Notebook to produce publication ready figures to be included in reports In Detail Python, a multi-paradigm programming language, has become the language of choice for data scientists for data analysis, visualization, and machine learning.

Ever imagined how to become an expert at effectively approaching data analysis problems, solving them, and extracting all of the available information from your data? Well, look no further, this is the book you want! Through this comprehensive guide, you will explore data and present results and conclusions from statistical analysis in a meaningful way. You’ll be able to quickly and accurately perform the hands-on sorting, reduction, and subsequent analysis, and fully appreciate how data analysis methods can support business decision-making. You’ll start off by learning about the tools available for data analysis in Python and will then explore the statistical models that are used to identify patterns in data. Gradually, you’ll move on to review statistical inference using Python, Pandas, and SciPy. After that, we’ll focus on performing regression using computational tools and you’ll get to understand the problem of identifying clusters in data in an algorithmic way. Finally, we delve into advanced techniques to quantify cause and effect using Bayesian methods and you’ll discover how to use Python’s tools for supervised machine learning. Style and approach This book takes a step-by-step approach to reading, processing, and analyzing data in Python using various methods and tools. Rich in examples, each topic connects to real-world examples and retrieves data directly online where possible. With this book, you are given the knowledge and tools to explore any data on your own, encouraging a curiosity befitting all data scientists.

Table of Contents
  1. Preface
  2. What this book covers
  3. What you need for this book
  4. Who this book is for
  5. Conventions
  6. Reader feedback
  7. Customer support
  8. Downloading the example code
  9. Downloading the color images of this book
  10. Errata
  11. Piracy
  12. Questions
  13. 1. Tools of the Trade
    1. Before you start
    2. Using the notebook interface
    3. Imports
    4. An example using the Pandas library
    5. Summary
  14. 2. Exploring Data
    1. The General Social Survey
    2. Obtaining the data
    3. Reading the data
    4. Univariate data
    5. Histograms
    6. Making things pretty
    7. Characterization
    8. Concept of statistical inference
    9. Numeric summaries and boxplots
    10. Relationships between variables – scatterplots
    11. Summary
  15. 3. Learning About Models
    1. Models and experiments
    2. The cumulative distribution function
    3. Working with distributions
    4. The probability density function
    5. Where do models come from?
    6. Multivariate distributions
    7. Summary
  16. 4. Regression
    1. Introducing linear regression
    2. Getting the dataset
    3. Testing with linear regression
    4. Multivariate regression
    5. Adding economic indicators
    6. Taking a step back
    7. Logistic regression
    8. Some notes
    9. Summary
  17. 5. Clustering
    1. Introduction to cluster finding
    2. Starting out simple – John Snow on cholera
    3. K-means clustering
    4. Suicide rate versus GDP versus absolute latitude
    5. Hierarchical clustering analysis
    6. Reading in and reducing the data
    7. Hierarchical cluster algorithm
    8. Summary
  18. 6. Bayesian Methods
    1. The Bayesian method
    2. Credible versus confidence intervals
    3. Bayes formula
    4. Python packages
    5. U.S. air travel safety record
    6. Getting the NTSB database
    7. Binning the data
    8. Bayesian analysis of the data
    9. Binning by month
    10. Plotting coordinates
    11. Cartopy
    12. Mpl toolkits – basemap
    13. Climate change – CO2 in the atmosphere
    14. Getting the data
    15. Creating and sampling the model
    16. Summary
  19. 7. Supervised and Unsupervised Learning
    1. Introduction to machine learning
    2. Scikit-learn
    3. Linear regression
    4. Climate data
    5. Checking with Bayesian analysis and OLS
    6. Clustering
    7. Seeds classification
    8. Visualizing the data
    9. Feature selection
    10. Classifying the data
    11. The SVC linear kernel
    12. The SVC Radial Basis Function
    13. The SVC polynomial
    14. K-Nearest Neighbour
    15. Random Forest
    16. Choosing your classifier
    17. Summary
  20. 8. Time Series Analysis
    1. Introduction
    2. Pandas and time series data
    3. Indexing and slicing
    4. Resampling, smoothing, and other estimates
    5. Stationarity
    6. Patterns and components
    7. Decomposing components
    8. Differencing
    9. Time series models
    10. Autoregressive – AR
    11. Moving average – MA
    12. Selecting p and q
    13. Automatic function
    14. The (Partial) AutoCorrelation Function
    15. Autoregressive Integrated Moving Average – ARIMA
    16. Summary
    17. A. More on Jupyter Notebook and matplotlib Styles
    18. Jupyter Notebook
    19. Useful keyboard shortcuts
    20. Command mode shortcuts
    21. Edit mode shortcuts
    22. Markdown cells
    23. Notebook Python extensions
    24. Installing the extensions
    25. Codefolding
    26. Collapsible headings
    27. Help panel
    28. Initialization cells
    29. NbExtensions menu item
    30. Ruler
    31. Skip-traceback
    32. Table of contents
    33. Other Jupyter Notebook tips
    34. External connections
    35. Export
    36. Additional file types
    37. Matplotlib styles
    38. Useful resources
    39. General resources
    40. Packages
    41. Data repositories
    42. Visualization of data
    43. Summary
Authors Biography

Magnus Vilhelm Persson is a scientist with a passion for Python and open source software usage and development. He obtained his PhD in Physics/Astronomy from Copenhagen University’s Centre for Star and Planet Formation (StarPlan) in 2013. Since then, he has continued his research in Astronomy at various academic institutes across Europe. In his research, he uses various types of data and analysis to gain insights into how stars are formed. He has participated in radio shows about Astronomy and also organized workshops and intensive courses about the use of Python for data analysis.

Luiz Felipe Martins holds a PhD in applied mathematics from Brown University and has worked as a researcher and educator for more than 20 years. His research is mainly in the field of applied probability. He has been involved in developing code for open source homework system, WeBWorK, where he wrote a library for the visualization of systems of differential equations. He was supported by an NSF grant for this project. Currently, he is an associate professor in the department of mathematics at Cleveland State University, Cleveland, Ohio, where he has developed several courses in applied mathematics and scientific computing. His current duties include coordinating all first-year calculus sessions.

Additional information
Weight0.489 kg
Reviews

There are no reviews yet.

Only logged in customers who have purchased this product may leave a review.