S3Bucket
A class for interacting with AWS S3 buckets using the boto3 library.
This class provides methods to read from and write to S3 buckets, including uploading and downloading files, listing bucket contents, and converting pandas DataFrames to CSV files for storage in S3.
Attributes:
| Name | Type | Description |
|---|---|---|
s3_client |
boto3.client
|
A boto3 client for interacting with AWS S3. |
Methods
read_csv_to_df(bucket_name, file_name): Reads a CSV file from an S3 bucket and returns it as a DataFrame. list_s3_bucket_contents(bucket_name): Lists all objects in an S3 bucket. write_df_to_csv(bucket_name, df, filename): Writes a DataFrame to a CSV file and uploads it to an S3 bucket. save_df_to_s3_if_not_exists(bucket_name, df, df_name): Saves a DataFrame to S3 as a CSV file if it doesn't exist. upload_binary_file(bucket_name, local_file_path, s3_key): Uploads a binary file to an S3 bucket. download_binary_file(bucket_name, s3_key, local_file_path): Downloads a binary file from an S3 bucket.
Example
s3 = S3Bucket('my_aws_profile') df = s3.read_csv_to_df('my_bucket', 'data.csv') s3.write_df_to_csv('my_bucket', df, 'data_backup.csv')
Note
Ensure that AWS credentials are properly configured for the boto3 client.
The class requires an AWS profile name to be passed during initialization for setting up the boto3 session.
Raises:
| Type | Description |
|---|---|
botocore.exceptions.ClientError
|
If operations on the S3 bucket fail due to client-side issues. |
botocore.exceptions.NoCredentialsError
|
If AWS credentials are not correctly configured or found. |
Source code in veetility/s3_bucket.py
8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 | |
__init__(aws_profile)
Initializes the S3Bucket class with the specified AWS profile.
This constructor sets up a new AWS session using the provided profile name and creates an S3 client for interacting with AWS S3 services. The S3 client is stored as an attribute for use in other methods of the class. This setup is essential for performing operations such as reading from and writing to S3 buckets.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
aws_profile |
str
|
The name of the AWS profile to use for creating the boto3 session. This profile should be configured in your AWS credentials file. |
required |
Raises:
| Type | Description |
|---|---|
boto3.exceptions.NoCredentialsError
|
If the specified AWS profile is not found or the credentials are invalid. |
botocore.exceptions.ClientError
|
If the AWS session or S3 client creation fails due to other AWS client side issues. |
Example
s3 = S3Bucket('my_aws_profile')
Note
Before using this class, ensure that the AWS credentials file is properly set up with the specified profile and that the profile has the necessary permissions to access AWS S3 services.
Source code in veetility/s3_bucket.py
38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 | |
download_binary_file(bucket_name, s3_key, local_file_path)
Download a binary file from an S3 bucket.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
s3_key |
str
|
S3 key for the file to download |
required |
local_file_path |
str
|
Path to the local destination |
required |
Returns:
| Type | Description |
|---|---|
None |
Source code in veetility/s3_bucket.py
247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 | |
download_s3_file(url, aws_access_key_id, aws_secret_access_key)
Download a file from an S3 Bucket given a url.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
url |
S3 URL |
required |
Returns:
| Type | Description |
|---|---|
File content as string |
Source code in veetility/s3_bucket.py
212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 | |
list_s3_bucket_contents(bucket_name, folder_path='')
Lists the contents of the specified folder in an S3 bucket.
This method retrieves a list of all objects (files) in the given folder of an S3 bucket and returns their keys (filenames). If the folder is empty, does not exist, or the specified bucket does not exist, it returns an empty list.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bucket_name |
str
|
The name of the S3 bucket. |
required |
folder_path |
str
|
The path of the folder within the S3 bucket. Defaults to the root of the bucket. |
''
|
Returns:
| Type | Description |
|---|---|
List[str]: A list of keys (filenames) of all objects in the specified folder of the S3 bucket. If the folder is empty, does not exist, or the bucket does not exist, an empty list is returned. |
Raises:
| Type | Description |
|---|---|
boto3.exceptions.S3Error
|
If an error occurs while accessing the S3 bucket. |
Examples:
>>> s3 = S3Bucket()
>>> s3.list_s3_bucket_contents('my_bucket', 'my_folder/')
['my_folder/file1.csv', 'my_folder/file2.csv', 'my_folder/image1.png']
Note
This method assumes that the AWS credentials and permissions are correctly set up to access the specified S3 bucket and folder.
Source code in veetility/s3_bucket.py
105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 | |
read_csv_to_df(bucket_name, folder_path='', file_name='')
Reads a CSV file from a specified folder in an AWS S3 bucket and converts it into a pandas DataFrame.
This function fetches the specified CSV file from the given folder in an S3 bucket and reads it into a pandas DataFrame. The CSV file is read into memory as a binary buffer using BytesIO, and then pandas is used to parse the CSV data from this buffer.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bucket_name |
str
|
The name of the S3 bucket from which the CSV file is read. |
required |
folder_path |
str
|
The path of the folder within the S3 bucket. Defaults to the root of the bucket. |
''
|
file_name |
str
|
The name (key) of the CSV file within the specified folder of the S3 bucket. |
''
|
Returns:
| Type | Description |
|---|---|
pandas.DataFrame: A DataFrame containing the data from the CSV file. |
Raises:
| Type | Description |
|---|---|
botocore.exceptions.ClientError
|
If the file is not found in the S3 bucket or the request to S3 fails. |
ValueError
|
If the |
Examples:
>>> s3 = S3Bucket()
>>> df = s3.read_csv_to_df('my_bucket', 'my_folder/', 'my_data.csv')
>>> print(df)
Note
Ensure that AWS credentials are properly configured and the boto3 client
is initialized before calling this function. This function assumes that the
CSV file is properly formatted for use with pandas' read_csv function.
Source code in veetility/s3_bucket.py
66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 | |
save_df_to_s3_if_not_exists(bucket_name, df, df_name, filetype, override=False)
Saves a pandas DataFrame to an S3 bucket as a CSV file if a file with the same name does not already exist.
This method checks the specified S3 bucket for an existing file that matches the naming convention '{df_name}_{current_date}.csv'. If such a file exists, it does not perform any action. Otherwise, it saves the provided DataFrame to the S3 bucket with the constructed filename.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bucket_name |
str
|
The name of the S3 bucket where the file will be saved. |
required |
df |
pandas.DataFrame
|
The DataFrame to be saved as a CSV file. |
required |
df_name |
str
|
The base name to be used for the CSV file. The current date in the format 'YYYY_MM_DD' will be appended to this base name. |
required |
Returns:
| Name | Type | Description |
|---|---|---|
None | This method does not return anything. It either saves the file to S3 or prints a message indicating that the file already exists. |
Raises:
| Type | Description |
|---|---|
boto3.exceptions.S3UploadFailedError
|
If the upload to the S3 bucket fails. |
Exception
|
If any other error occurs during the process. |
Examples:
>>> s3 = S3Bucket()
>>> s3.save_df_to_s3_if_not_exists('my_bucket', my_dataframe, 'sales_data')
File 'sales_data_2023_11_22.csv' successfully uploaded to bucket 'my_bucket'.
Source code in veetility/s3_bucket.py
170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 | |
upload_binary_file(bucket_name, local_file_path, s3_key)
Upload a binary file to an S3 bucket.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
local_file_path |
str
|
Path to the local file |
required |
s3_key |
str
|
S3 key for the uploaded file |
required |
Returns:
| Type | Description |
|---|---|
None |
Source code in veetility/s3_bucket.py
229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 | |
write_df_to_file(bucket_name, df, filename, filetype='csv')
Write a DataFrame to a file and upload it to an S3 bucket.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
bucket_name |
str
|
Name of the S3 bucket. |
required |
df |
pandas.DataFrame
|
DataFrame to be written to file. |
required |
filename |
str
|
The name of the file to be created. |
required |
filetype |
str
|
Type of file to create ('csv' or 'json'). Defaults to 'csv'. |
'csv'
|
Raises:
| Type | Description |
|---|---|
ValueError
|
If the specified filetype is not supported. |
Source code in veetility/s3_bucket.py
140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 | |