I wanted to learn more about AWS because as Amazon says, they want to "enable developers to focus on innovating with data, rather than how to store it". You don't have the pains of being required to design the infrastructure needs of the system as a whole now in the beginning, or the future when you need performance and reliability. It has 99.999999999% durability, with 99.99% availability by replicating across several facilities in a region. It scales automatically, you don't do anything, and it remains available as it's doing that under the covers. Amazon has a great 'Free Usage Tier' for folks like me (and you) who are just starting out and want to get hands-on experience. Not all of the services are offered in the free usage tier (S3 and DynamoDB are), so take a look!
Summaries
Security credentials
There is a AWS root account credential. One of the first things you should do is create a user for yourself, and assign it to the admin group. Never use your AWS root account credential directly! You can then create IAM (Identity and Access Management) users depending on the application need, creating groups with logical functions, and users to those groups.
S3
- Literally don't have to configure anything to get started, you have a key, and you upload your value, just keep storing things in S3. You would still need some way to keep track of what keys you are using for you specific application
- Doesn't support object locking, you'll need to build this manually if it's needed
- S3 max data storage is 5TB, no limit on attributes for an item
- Uses the eventual consistency consistency model
DynamoDB
- Minimum to configure when creating your tables is to specify the table primary key, and provisioning the read and write throughput needed
- Supports optimistic object locking
- Max data size is 64KB, no limit on number of attributes for an item
- Supports eventual consistency or strongly consistent consistency model
When to use S3 versus DynamoDB?
S3 is for larger size data that you are rarely going to touch (like storing data for analysis or backup/archiving of data). Whereas DynamoDB is more for data that is more dynamic. As we already talked about with both, you can start with a small amount of data, and scale up and down on the requirements. For DynamoDB, you would also need to adjust your read and write capacities. In your custom application, you can also use a mix both, or use other Amazon services simultaneously.
The storage for S3 and DynamoDB are relatively cheap, and they only get cheaper. So if you are developing an app for your local, dev, test, or pre-prod environments, you just may want to go ahead and use different instances of the services. Or you can use a sandbox tool like ZCloud AWS sandbox for S3.
Get the AWS SDK for Java:
- For java / maven folks, the easiest way is to use the maven dependency to grab the AWS SDK for java at http://mvnrepository.com/artifact/com.amazonaws/aws-java-sdk
- For eclipse folks, the easiest way is to get the AWS Toolkit plugin
- S3 sample project to get started provided by amazon
- DynamoDB sample code to get started provided by amazon
Next in the line on NoSQL database are Apache Cassandra and MongoDB, I'll have more thoughts on those in the coming months.
For more examples, ideas, and inspiration, feel free to read through the links provided in the "Now what, want to start playing around" section. What questions do you have about this post? Let me know in the comments section below, and I will answer each one.
********
For details on the code that I played around with, I grabbed AWS SDK for Java with Maven, and did everything in JUnit integration tests. Starting point for most of this code was from the sample code links from above.