Python Data Analysis (2nd Edition)

Rs. 13,975
  • Author: Armando Fandango
  • ISBN: 9781787127487
  • Publisher: Packt Publishing
  • Edition: 2nd
  • Publication Date: March 30, 2017
  • Format: Paperback – 330 pages
  • Language: English

Out of stock



Share
Description
Key Features
  • Find, manipulate, and analyze your data using the Python 3.5 libraries
  • Perform advanced, high-performance linear algebra and mathematical calculations with clean and efficient Python code
  • An easy-to-follow guide with realistic examples that are frequently used in real-world data analysis projects.

Who This Book Is For

This book is for programmers, scientists, and engineers who have the knowledge of Python and know the basics of data science. It is for those who wish to learn different data analysis methods using Python 3.5 and its libraries. This book contains all the basic ingredients you need to become an expert data analyst.

What You Will Learn

  • Install open source Python modules such NumPy, SciPy, Pandas, stasmodels, scikit-learn,theano, keras, and tensorflow on various platforms
  • Prepare and clean your data, and use it for exploratory analysis
  • Manipulate your data with Pandas
  • Retrieve and store your data from RDBMS, NoSQL, and distributed filesystems such as HDFS and HDF5
  • Visualize your data with open source libraries such as matplotlib, bokeh, and plotly
  • Learn about various machine learning methods such as supervised, unsupervised, probabilistic, and Bayesian
  • Understand signal processing and time series data analysis
  • Get to grips with graph processing and social network analysis

In Detail

Data analysis techniques generate useful insights from small and large volumes of data. Python, with its strong set of libraries, has become a popular platform to conduct various data analysis and predictive modeling tasks.

With this book, you will learn how to process and manipulate data with Python for complex analysis and modeling. We learn data manipulations such as aggregating, concatenating, appending, cleaning, and handling missing values, with NumPy and Pandas. The book covers how to store and retrieve data from various data sources such as SQL and NoSQL, CSV fies, and HDF5. We learn how to visualize data using visualization libraries, along with advanced topics such as signal processing, time series, textual data analysis, machine learning, and social media analysis.

The book covers a plethora of Python modules, such as matplotlib, statsmodels, scikit-learn, and NLTK. It also covers using Python with external environments such as R, Fortran, C/C++, and Boost libraries.

Style and approach

The book takes a very comprehensive approach to enhance your understanding of data analysis. Sufficient real-world examples and use cases are included in the book to help you grasp the concepts quickly and apply them easily in your day-to-day work. Packed with clear, easy to follow examples, this book will turn you into an ace data analyst in no time.

Table of Contents
  1. Python Data Analysis – Second Edition
  2. Credits
  3. About the Author
  4. About the Reviewers
  5. Why subscribe?
  6. Customer Feedback
  7. Preface
  8. What this book covers
  9. What you need for this book
  10. Who this book is for
  11. Conventions
  12. Reader feedback
  13. Customer support
  14. Downloading the example code
  15. Downloading the color images of this book
  16. Errata
  17. Piracy
  18. Questions
  19. 1. Getting Started with Python Libraries
  20. Installing Python 3
  21. Installing data analysis libraries
  22. On Linux or Mac OS X
  23. On Windows
  24. Using IPython as a shell
  25. Reading manual pages
  26. Jupyter Notebook
  27. NumPy arrays
  28. A simple application
  29. Where to find help and references
  30. Listing modules inside the Python libraries
  31. Visualizing data using Matplotlib
  32. Summary
  33. 2. NumPy Arrays
  34. The NumPy array object
  35. Advantages of NumPy arrays
  36. Creating a multidimensional array
  37. Selecting NumPy array elements
  38. NumPy numerical types
  39. Data type objects
  40. Character codes
  41. The dtype constructors
  42. The dtype attributes
  43. One-dimensional slicing and indexing
  44. Manipulating array shapes
  45. Stacking arrays
  46. Splitting NumPy arrays
  47. NumPy array attributes
  48. Converting arrays
  49. Creating array views and copies
  50. Fancy indexing
  51. Indexing with a list of locations
  52. Indexing NumPy arrays with Booleans
  53. Broadcasting NumPy arrays
  54. Summary
  55. References
  56. 3. The Pandas Primer
  57. Installing and exploring Pandas
  58. The Pandas DataFrames
  59. The Pandas Series
  60. Querying data in Pandas
  61. Statistics with Pandas DataFrames
  62. Data aggregation with Pandas DataFrames
  63. Concatenating and appending DataFrames
  64. Joining DataFrames
  65. Handling missing values
  66. Dealing with dates
  67. Pivot tables
  68. Summary
  69. References
  70. 4. Statistics and Linear Algebra
  71. Basic descriptive statistics with NumPy
  72. Linear algebra with NumPy
  73. Inverting matrices with NumPy
  74. Solving linear systems with NumPy
  75. Finding eigenvalues and eigenvectors with NumPy
  76. NumPy random numbers
  77. Gambling with the binomial distribution
  78. Sampling the normal distribution
  79. Performing a normality test with SciPy
  80. Creating a NumPy masked array
  81. Disregarding negative and extreme values
  82. Summary
  83. 5. Retrieving, Processing, and Storing Data
  84. Writing CSV files with NumPy and Pandas
  85. The binary .npy and pickle formats
  86. Storing data with PyTables
  87. Reading and writing Pandas DataFrames to HDF5 stores
  88. Reading and writing to Excel with Pandas
  89. Using REST web services and JSON
  90. Reading and writing JSON with Pandas
  91. Parsing RSS and Atom feeds
  92. Parsing HTML with Beautiful Soup
  93. Summary
  94. Reference
  95. 6. Data Visualization
  96. The matplotlib subpackages
  97. Basic matplotlib plots
  98. Logarithmic plots
  99. Scatter plots
  100. Legends and annotations
  101. Three-dimensional plots
  102. Plotting in Pandas
  103. Lag plots
  104. Autocorrelation plots
  105. Plot.ly
  106. Summary
  107. 7. Signal Processing and Time Series
  108. The statsmodels modules
  109. Moving averages
  110. Window functions
  111. Defining cointegration
  112. Autocorrelation
  113. Autoregressive models
  114. ARMA models
  115. Generating periodic signals
  116. Fourier analysis
  117. Spectral analysis
  118. Filtering
  119. Summary
  120. 8. Working with Databases
  121. Lightweight access with sqlite3
  122. Accessing databases from Pandas
  123. SQLAlchemy
  124. Installing and setting up SQLAlchemy
  125. Populating a database with SQLAlchemy
  126. Querying the database with SQLAlchemy
  127. Pony ORM
  128. Dataset – databases for lazy people
  129. PyMongo and MongoDB
  130. Storing data in Redis
  131. Storing data in memcache
  132. Apache Cassandra
  133. Summary
  134. 9. Analyzing Textual Data and Social Media
  135. Installing NLTK
  136. About NLTK
  137. Filtering out stopwords, names, and numbers
  138. The bag-of-words model
  139. Analyzing word frequencies
  140. Naive Bayes classification
  141. Sentiment analysis
  142. Creating word clouds
  143. Social network analysis
  144. Summary
  145. 10. Predictive Analytics and Machine Learning
  146. Preprocessing
  147. Classification with logistic regression
  148. Classification with support vector machines
  149. Regression with ElasticNetCV
  150. Support vector regression
  151. Clustering with affinity propagation
  152. Mean shift
  153. Genetic algorithms
  154. Neural networks
  155. Decision trees
  156. Summary
  157. 11. Environments Outside the Python Ecosystem and Cloud Computing
  158. Exchanging information with Matlab/Octave
  159. Installing rpy2 package
  160. Interfacing with R
  161. Sending NumPy arrays to Java
  162. Integrating SWIG and NumPy
  163. Integrating Boost and Python
  164. Using Fortran code through f2py
  165. PythonAnywhere Cloud
  166. Summary
  167. 12. Performance Tuning, Profiling, and Concurrency
  168. Profiling the code
  169. Installing Cython
  170. Calling C code
  171. Creating a process pool with multiprocessing
  172. Speeding up embarrassingly parallel for loops with Joblib
  173. Comparing Bottleneck to NumPy functions
  174. Performing MapReduce with Jug
  175. Installing MPI for Python
  176. IPython Parallel
  177. Summary
  178. A. Key Concepts
  179. B. Useful Functions
  180. Matplotlib
  181. NumPy
  182. Pandas
  183. Scikit-learn
  184. SciPy
  185. scipy.fftpack
  186. scipy.signal
  187. scipy.stats
  188. C. Online Resources
Author Biography

Armando Fandango is Chief Data Scientist at Epic Engineering and Consulting Group, and works on confidential projects related to defense and government agencies. Armando is an accomplished technologist with hands-on capabilities and senior executive-level experience with startups and large companies globally. His work spans diverse industries including FinTech, stock exchanges, banking, bioinformatics, genomics, AdTech, infrastructure, transportation, energy, human resources, and entertainment.
Armando has worked for more than ten years in projects involving predictive analytics, data science, machine learning, big data, product engineering, high performance computing, and cloud infrastructures. His research interests spans machine learning, deep learning, and scientific computing.

Additional information
Weight0.566 kg
Reviews

There are no reviews yet.

Only logged in customers who have purchased this product may leave a review.