Friday, January 29, 2016

Coursera Regression Course - Assignment 2

A major life depression is associated with
a higher daily consumption of alcohol

All data from the NESARC Study

Categorical Explanatory Variable

We have a categorical variable encoded with No=0, implying that subjects with no major depression in life are at the baseline 0.
3549-3549
MAJORDEPLIFE
MAJOR DEPRESSION - LIFETIME (NON-HIERARCHICAL)
----------------------------------------------
35254 0. No
7839 1. Yes
----------------------------------------------

Quantitative Response Variable

The amount of ethanol (alcohol) consumed averagely in the past year is our response variable.
3675-3682
ETOTLCA2
AVERAGE DAILY VOLUME OF ETHANOL CONSUMED IN PAST YEAR, FROM ALL TYPES OF ALCOHOLIC BEVERAGES COMBINED
(NOTE: Users may wish to exclude outliers)
--------------------------------------------------
0.0003 - 219.9555 Ounces of ethanol
Blank Unknown
--------------------------------------------------

We see that there is a large number of abstinent people in the dataset. We still don't know if they are associated with depression or not. It would be interesting to see if they are more the one or the other.

Linear Regression Test

                            OLS Regression Results                            
==============================================================================
Dep. Variable:               ETOTLCA2   R-squared:                       0.001
Model:                            OLS   Adj. R-squared:                  0.001
Method:                 Least Squares   F-statistic:                     18.86
Date:                Fri, 29 Jan 2016   Prob (F-statistic):           1.42e-05
Time:                        15:57:58   Log-Likelihood:                -59022.
No. Observations:               26655   AIC:                         1.180e+05
Df Residuals:                   26653   BIC:                         1.181e+05
Df Model:                           1                                         
Covariance Type:            nonrobust                                         
================================================================================
                   coef    std err          t      P>|t|      [95.0% Conf. Int.]
--------------------------------------------------------------------------------
Intercept        0.5366      0.015     35.377      0.000         0.507     0.566
MAJORDEPLIFE     0.1474      0.034      4.342      0.000         0.081     0.214
==============================================================================
Omnibus:                    84158.156   Durbin-Watson:                   2.010
Prob(Omnibus):                  0.000   Jarque-Bera (JB):      20279387185.774
Skew:                          50.216   Prob(JB):                         0.00
Kurtosis:                    4274.926   Cond. No.                         2.62
==============================================================================

The R-squared number is rather low, sugesting that only 0,1% of the "Etanolic Intake" can be explained by a "major depression in life". At least the F-statistic is so low that we can assume the small effect is statistically significant. 
Now I would not go further but lets assume that R-squared is higher, in that case we believe that a major depression in a lifetime can predict Beta=14.74% increase in alcohol intake, with a high statistical significance of P>0.000

Graph: a major life depression could increase average ethanol intake from 54% to 68%

Program Code and Output

In [1]:
%matplotlib inline

import numpy
import pandas
import statsmodels.api as sm
import seaborn
import matplotlib.pyplot as plt
import statsmodels.formula.api as smf 
In [2]:
# bug fix for display formats to avoid run time errors
pandas.set_option('display.float_format', lambda x:'%.2f'%x)
In [3]:
data = pandas.read_csv('nesarc_pds.csv', low_memory=False)
In [4]:
#setting variables you will be working with to numeric
data['p'] = pandas.to_numeric(data['MAJORDEPLIFE'], errors='coerce')
In [5]:
data['ETOTLCA2'] = pandas.to_numeric(data['ETOTLCA2'], errors='coerce')
In [6]:
subGroupEthanol = list(filter(lambda x: x > 10.0, data['ETOTLCA2'].dropna()))
seaborn.distplot(subGroupEthanol)
Out[6]:
In [7]:
reg1 = smf.ols('ETOTLCA2 ~ MAJORDEPLIFE', data=data).fit()
print (reg1.summary())
                            OLS Regression Results                            
==============================================================================
Dep. Variable:               ETOTLCA2   R-squared:                       0.001
Model:                            OLS   Adj. R-squared:                  0.001
Method:                 Least Squares   F-statistic:                     18.86
Date:                Fri, 29 Jan 2016   Prob (F-statistic):           1.42e-05
Time:                        17:11:29   Log-Likelihood:                -59022.
No. Observations:               26655   AIC:                         1.180e+05
Df Residuals:                   26653   BIC:                         1.181e+05
Df Model:                           1                                         
Covariance Type:            nonrobust                                         
================================================================================
                   coef    std err          t      P>|t|      [95.0% Conf. Int.]
--------------------------------------------------------------------------------
Intercept        0.5366      0.015     35.377      0.000         0.507     0.566
MAJORDEPLIFE     0.1474      0.034      4.342      0.000         0.081     0.214
==============================================================================
Omnibus:                    84158.156   Durbin-Watson:                   2.010
Prob(Omnibus):                  0.000   Jarque-Bera (JB):      20279387185.774
Skew:                          50.216   Prob(JB):                         0.00
Kurtosis:                    4274.926   Cond. No.                         2.62
==============================================================================

Warnings:
[1] Standard Errors assume that the covariance matrix of the errors is correctly specified.
In [8]:
# listwise deletion for calculating means for regression model observations
sub1 = data[['ETOTLCA2', 'MAJORDEPLIFE']].dropna()

# group means & sd
print ("Mean")
ds1 = sub1.groupby('MAJORDEPLIFE').mean()
print (ds1)
print ("Standard deviation")
ds2 = sub1.groupby('MAJORDEPLIFE').std()
print (ds2)

# bivariate bar graph
seaborn.factorplot(x="MAJORDEPLIFE", y="ETOTLCA2", data=sub1, kind="bar", ci=None)
plt.xlabel('Major Life Depression')
plt.ylabel('Mean Number AVERAGE DAILY VOLUME OF ETHANOL CONSUMED IN PAST YEAR')
Mean
              ETOTLCA2
MAJORDEPLIFE          
0                 0.54
1                 0.68
Standard deviation
              ETOTLCA2
MAJORDEPLIFE          
0                 1.44
1                 4.03
Out[8]:
In [ ]:
 

Wednesday, January 20, 2016

First Assignment Coursera - Regression Models in Practice


I have chosen the National Epidemiologic Survey on Alcohol and Related Conditions (NESARC)

http://pubs.niaaa.nih.gov/publications/AA70/AA70.htm


1. Describe the sample

a) Describe the study population (who or what was studied).

NESARC contains an extensive battery of questions about present and past alcohol consumption, AUDs, and the use of alcohol treatment services. NESARC also included similar questions related to tobacco and illicit drug use (including nicotine dependence and drug use disorders) as well as questions designed to determine a wide variety of psychiatric disorders such as major depression, anxiety disorders, and personality disorders. 
 Furthermore, NESARC contained questions that operationalized the criteria set forth in the American Psychiatric Association’s Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition (DSM–IV) for the following psychiatric disorders:
  • Five mood disorders (major depressive disorder, bipolar I and bipolar II disorders, dysthymia, and hypomania)
  • Four anxiety disorders (panic with and without agoraphobia, social phobia, specific phobia, and generalized anxiety)
  • Seven personality disorders (avoidant, dependent, obsessive–compulsive, paranoid, schizoid, histrionic, and antisocial disorders).

b) Report the level of analysis studied (individual, group, or aggregate).

Individuals were randomly selected from a systematic sample of group quarters in these PSUs.The Census 2000 Group Quarters Inventory formed the sampling frame for the group quarters portion of the NESARC sample. 

c) Report the number of observations in the data set.

43093 Total, 8209 1. Northeast 8991 2. Midwest 16156 3. South 9737 4. West
d) Describe your data analytic sample (the sample you are using for your analyses).
397-398 S2AQ8A HOW OFTEN DRANK ANY ALCOHOL IN LAST 12 MONTHS
 --------------------------------------------- 
 1865 1. Every day 
 1210 2. Nearly every day 
 2619 3. 3 to 4 times a week 
 2914 4. 2 times a week 
 3261 5. Once a week 
 3557 6. 2 to 3 times a month 
 2663 7. Once a month 1805 8. 7 to 11 times in the last year 
 3210 9. 3 to 6 times in the last year 
 3637 10. 1 or 2 times in the last year 
 205 99. Unknown 
16147 BL. NA, former drinker or lifetime abstainer 
---------------------------------------------------------------------------------------

2. Describe the data collection procedure

a) Report the study design that generated that data (for example: data reporting, surveys, observation, experiment).

Surveys, Data were collected in face-to-face, computer-assisted personal interviews conducted in respondents’ homes. The NESARC response rate was 81 percent.

b) Describe the original purpose of the data collection.

The major purposes of the Wave 1 and Wave 2 NESARC are to:
  • Determine the prevalence, incidence, stability, and recurrence of AUDs and their associated disabilities in the general U.S. population.
  • Estimate the magnitude of health disparities in AUDs and their associated disabilities among population subgroups defined by gender, race/ethnicity, disability, sexual orientation, age, and socioeconomic status.
  • Estimate the size, characteristics, and changing nature of populations of special concern, including alcohol abusers and other people in the general population who are impaired or affected by the use of alcohol (e.g., those engaging in binge drinking or impaired driving).
  • Estimate changes in AUDs and their associated disabilities over time, and identify factors associated with the natural history of AUDs.
  • Determine the number of people receiving alcohol treatment through various treatment programs and services, including those not otherwise represented in surveys of treatment facilities; measure the unmet need for alcohol treatment services; and identify barriers to seeking treatment.
  • Determine the associations between AUDs and their major physical and mental disabilities, differentiating drug-induced disorders from those reflecting true, independent mental conditions.
  • Determine the boundaries between safe and hazardous drinking levels and patterns for various types of AUDs and their associated medical, social, and psychological sequelae.

c) Describe how the data were collected.

NESARC participants are variable in lifestyle and ages.

To ensure that minority and special populations were well represented in the sample, NESARC oversampled Blacks, Hispanics, and young adults ages 18–24. As a result, the survey produced enough minority respondents to answer questions of race/ethnic disparities in comorbidity and access to health care services

d) Report when the data were collected.

In 2001—2002, the National Institute on Alcohol Abuse and Alcoholism (NIAAA) conducted the first wave of the National Epidemiologic Survey on Alcohol and Related Conditions (NESARC),The second wave of the survey took place from 2004 to 2005. 

e) Report where the data were collected.

They represented all regions of the United States. In addition to sampling people living in traditional households, NESARC investigators questioned military personnel living off base and people living in a variety of group accommodations such as boarding or rooming houses and college quarters. By including these different types of housing, the investigators were able to obtain data on people not typically captured by household surveys.

 3. Measures section describing your variables and how you managed them to address your own research question

a) Describe what your explanatory and response variables measured.

explanatory variable:

397-398 S2AQ8A HOW OFTEN DRANK ANY ALCOHOL IN LAST 12 MONTHS
 --------------------------------------------- 
 1865 1. Every day 
 1210 2. Nearly every day 
 2619 3. 3 to 4 times a week 
 2914 4. 2 times a week 
 3261 5. Once a week 
 3557 6. 2 to 3 times a month 
 2663 7. Once a month 1805 8. 7 to 11 times in the last year 
 3210 9. 3 to 6 times in the last year 
 3637 10. 1 or 2 times in the last year 
 205 99. Unknown 
16147 BL. NA, former drinker or lifetime abstainer 
---------------------------------------------------------------------------------------
b) Describe the response scales for your explanatory and response variables.
c) Describe how you managed your explanatory and response variables.

response variable

 SECTION 9: GENERALIZED ANXIETY (GENERAL ANXIETY) -
 3020-3020 S9Q1A EVER HAD 6+ MONTH PERIOD FELT TENSE/NERVOUS/WORRIED MOST OF TIME
 3128 1. Yes 38225 2. No 1740 9. Unknown

 3021-3021 S9Q1B EVER HAD 6+ MONTH PERIOD FELT VERY TENSE/NERVOUS/WORRIED MOST OF TIME ABOUT EVERYDAY PROBLEMS
 358 1. Yes 37867 2. No 1740 9. Unknown 3128 BL. NA, had 6+ month period feeling tense/nervous/worried most of the time

 3022-3022 S9Q31 IN WORST PERIOD, EVER WORRY A LOT ABOUT THINGS YOU USUALLY DIDN'T WORRY ABOUT
2336 1. Yes 1115 2. No 35 9. Unknown 39607 BL. NA, never or unkn. if ever had 6+ month pd. being tense/nervous/worried

 3023-3023 S9Q32 IN WORST PERIOD, EVER WORRY ABOUT MORE THAN ONE THING
 2644 1. Yes 808 2. No 34 9. Unknown 39607 BL. NA, never or unkn. if ever had 6+ month pd. being tense/nervous/worried

 3024-3024 S9Q33 IN WORST PERIOD, EVER FIND IT DIFFICULT TO STOP BEING TENSE/NERVOUS/WORRIED
 2715 1. Yes 742 2. No 29 9. Unknown 39607 BL. NA, never or unkn. if ever had 6+ month pd. being tense/nervous/worried

 3025-3025 S9Q34 IN WORST PERIOD, EVER WORRY ABOUT THINGS THAT WERE VERY UNLIKELY TO HAPPEN
1418 1. Yes 2019 2. No 49 9. Unknown 39607 BL. NA, never or unkn. if ever had 6+ month pd. being tense/nervous/worried

c) Describe how you managed your explanatory and response variables.
confounders:

I will control for different confounders like:

 Page 302: SECTION 4A: MAJOR DEPRESSION (LOW MOOD I)
Page 316: SECTION 4B: FAMILY HISTORY (III) OF MAJOR DEPRESSION
Page 319: SECTION 4C: DYSTHYMIA (LOW MOOD II)
Page 329: SECTION 5: MANIA OR HYPOMANIA (HIGH MOOD)
Page 341: SECTION 6: PANIC DISORDERS AND AGORAPHOBIA (ANXIETY)
Page 355: SECTION 7: SOCIAL PHOBIA (SOCIAL SITUATIONS)
Page 371: SECTION 8: SPECIFIC PHOBIA (SPECIFIC SITUATIONS)
Page 386: SECTION 9: GENERALIZED ANXIETY (GENERAL ANXIETY)
Page 403: SECTION 10: PERSONALITY DISORDERS (USUAL FEELINGS/ACTIONS)
Page 419: SECTION 11A: ANTISOCIAL PERSONALITY DISORDER (BEHAVIOR)

Sunday, September 1, 2013

Which Travel Guide for Tyrol? To Germans does not matter

My endevor to build an tourism indicator with online booksales led to a first interesting result.
In the category of worldwide travelguides, you see that the rank number (more is worse) increases comparably fast.
Whereas within the Tyrol category there are alot of books doing comparably well. The rank number increases slow.

World travel guides 
Italy travel guides
South Tyrol travel guides
Tyrol travel guides

Monday, January 9, 2012

See you at G+

As long as I have not a longer article to publish you find daily rants and ramblings on certainty, uncertainty and ignorance under https://plus.google.com/u/0/103983654382933904290/posts