Getting Started
Hello World & Comments
Python uses # for comments and triple quotes for docstrings. The print() function supports sep and end parameters to customize output formatting. Docstrings serve as documentation accessible via help() and __doc__.
# This is a single-line comment
"""
This is a
multi-line comment (docstring)
"""
print("Hello, World!") # print to stdout
print("A", "B", "C", sep="-") # A-B-C
print("No newline", end="") # suppress newlineIndentation & Code Blocks
Unlike most languages, Python uses indentation rather than braces to define code blocks. Consistency is critical — mixing tabs and spaces causes SyntaxError. PEP 8 recommends 4 spaces per level.
# Python uses indentation (4 spaces) to define blocks
if True:
print("inside if")
if True:
print("nested block")
print("outside block")
# No braces! Indentation IS the syntax
def func():
x = 1
return x + 1Input & Output
input() reads from stdin as a string — always convert when you need a number. Use int(), float(), etc. for conversion. F-strings (Python 3.6+) are the preferred way to format strings.
# input() always returns a string
name = input("Enter your name: ")
age = int(input("Enter your age: ")) # convert to int
print(f"Hello {name}, you are {age} years old")
# formatted output
print("Pi is approximately {:.2f}".format(3.14159))
print(f"{1000000:,}") # 1,000,000 with thousands separatorMultiple Statements & Line Continuation
Use semicolons to separate statements on one line (rare in idiomatic Python). Long lines can be continued with backslash, or automatically inside (), [], {}. Prefer implicit continuation for readability.
# multiple statements on one line (discouraged)
a = 1; b = 2; c = 3
# explicit line continuation
total = 1 + 2 + 3 + \
4 + 5 + 6
# implicit continuation inside brackets
nums = [
1, 2, 3,
4, 5, 6
]
result = (1 + 2
+ 3 + 4)Python Execution
Python scripts run with 'python script.py'. The REPL allows interactive experimentation. Always use python3 on systems where python points to Python 2. The shebang line makes scripts executable on Unix.
# Run a script
# $ python script.py
# Run interactively (REPL)
# $ python
# >>> 2 + 2
# 4
# Shebang line for Unix scripts
#!/usr/bin/env python3
# Check Python version
import sys
print(sys.version)
print(sys.version_info.major) # 3Variables & Data Types
Variables & Dynamic Typing
Python uses dynamic typing — variables can change type at runtime. Use type() to check, isinstance() to verify. Python 3.6+ supports type hints (name: str = 'Alice') for IDE support without runtime enforcement.
# Python is dynamically typed - no declaration needed
name = "Alice" # str
age = 30 # int
height = 5.7 # float
is_active = True # bool
items = [1, 2, 3] # list
# Type checking
print(type(name)) # <class 'str'>
print(isinstance(age, int)) # True
# Multiple assignment
x, y, z = 1, 2, 3
a = b = 0 # chain assignmentType Hints (Python 3.6+)
Type hints improve code readability and enable IDE autocompletion and static analysis with mypy. They are NOT enforced at runtime — Python remains dynamically typed. Use Optional[X] for values that can be None.
# Variable annotations
name: str = "Alice"
age: int = 30
scores: list[float] = [90.5, 85.0]
# Function annotations
def greet(name: str, times: int = 1) -> str:
return (f"Hi {name}! " * times).strip()
# Optional and Union
from typing import Optional, Union
def find(id: int) -> Optional[str]:
return "Alice" if id == 1 else None
# mypy for static type checking
# $ mypy script.pyType Conversion
Python has built-in conversion functions: int(), float(), str(), bool(), list(), tuple(), set(), dict(). Falsy values include 0, '', [], {}, None, False. int() truncates toward zero, while round() uses banker's rounding.
# String to number
num_str = str(42) # "42"
num = int("42") # 42
float_num = float("3.14") # 3.14
# Number conversions
print(int(3.99)) # 3 (truncates toward zero)
print(int(-3.99)) # -3
print(round(3.14159, 2)) # 3.14
# Boolean conversion
print(bool(0)) # False
print(bool("")) # False
print(bool([])) # False
print(bool("anything")) # True
# Collection conversions
print(list("abc")) # ['a', 'b', 'c']
print(tuple([1, 2, 3])) # (1, 2, 3)
print(set([1, 1, 2])) # {1, 2}Numeric Types
Python ints have arbitrary precision (no overflow). Floats are IEEE 754 doubles with the usual precision issues. Complex numbers are built-in. Booleans are a subclass of int (True==1, False==0). Use underscores in numeric literals for readability.
# Integers (arbitrary precision)
big = 10 ** 100 # no overflow
print(type(big)) # <class 'int'>
# Floats (IEEE 754 double)
pi = 3.14159
print(0.1 + 0.2) # 0.30000000000000004
# Complex numbers
z = 3 + 4j
print(z.real, z.imag) # 3.0 4.0
print(abs(z)) # 5.0
# Boolean is subclass of int
print(isinstance(True, int)) # True
print(True + True) # 2
# Underscores in numbers (3.6+)
million = 1_000_000
binary = 0b_1010_1010Constants & Naming Conventions
Python has no const keyword — ALL_CAPS names are constants by convention only (nothing prevents reassignment). PEP 8 defines naming: snake_case for variables/functions, PascalCase for classes, ALL_CAPS for constants. Leading underscore means 'private' by convention; double underscore triggers name mangling.
# Python has no true constants - convention only
MAX_SIZE = 100 # ALL_CAPS for constants
PI = 3.14159
# Naming conventions (PEP 8)
variable_name = "snake_case" # variables, functions
ClassName = "PascalCase" # classes
CONSTANT_VALUE = 100 # constants
_private_var = "underscore prefix" # private (convention)
__name_mangled = "double underscore" # name mangling
# dunder names (reserved)
__name__, __main__, __init__Strings
String Methods
Strings are immutable — methods return new strings. find() returns -1 if not found, while index() raises ValueError. Use isalpha()/isdigit()/isalnum() for validation. The str class has over 40 methods — explore with dir(str).
s = "Hello, World"
# Case operations
print(s.upper()) # HELLO, WORLD
print(s.lower()) # hello, world
print(s.title()) # Hello, World
print(s.capitalize()) # Hello, world
print(s.swapcase()) # hELLO, wORLD
# Search & replace
print(s.find("World")) # 7 (index, -1 if not found)
print(s.index("World")) # 7 (raises ValueError if not found)
print(s.replace("o", "0")) # Hell0, W0rld
print(s.count("l")) # 3
# Validation
print("abc".isalpha()) # True
print("123".isdigit()) # True
print(" ".isspace()) # TrueString Formatting
F-strings are the modern, fastest, and most readable way to format strings. They support format specs after a colon: :.2f for 2 decimals, :>10 for right-align width 10, :, for thousands separator. Avoid %-formatting in new code.
name = "Alice"
age = 30
# f-strings (Python 3.6+) - PREFERRED
print(f"Hello, {name}! You are {age}.")
print(f"{name.upper()} is {age * 365} days old")
print(f"{3.14159:.2f}") # 3.14
print(f"{42:>10}") # right-align
print(f"{42:<10}") # left-align
print(f"{42:^10}") # center
print(f"{1000000:,}") # 1,000,000
# str.format() method
print("Hello, {}!".format(name))
print("{name} is {age}".format(name="Bob", age=25))
# Old style (avoid in new code)
print("Hello, %s!" % name)Slicing & Indexing
Python slicing syntax [start:stop:step] is powerful — stop is exclusive. Negative indices count from the end. s[::-1] is the idiomatic way to reverse a string. Slicing is safe: out-of-range indices return empty strings rather than raising errors.
s = "Hello, World"
# Indexing (0-based, negative from end)
print(s[0]) # H
print(s[-1]) # d
print(s[7]) # W
# Slicing [start:stop:step]
print(s[0:5]) # Hello
print(s[7:]) # World
print(s[:5]) # Hello
print(s[::2]) # HloWrd (every 2nd char)
print(s[::-1]) # dlroW ,olleH (reverse!)
# Length
print(len(s)) # 12
# Slicing never raises IndexError
print(s[100:200]) # '' (empty string)Splitting & Joining
split() divides a string into a list, join() combines a list into a string. Always use join() for efficient concatenation of many strings — the + operator creates intermediate strings. partition() splits into exactly 3 parts (before, separator, after).
# Split
csv = "a,b,c,d"
print(csv.split(",")) # ['a', 'b', 'c', 'd']
print(csv.split(",", 2)) # ['a', 'b', 'c,d'] (max 2 splits)
# Splitlines
text = "line1\nline2\nline3"
print(text.splitlines()) # ['line1', 'line2', 'line3']
# Partition (splits on first occurrence)
print("[email protected]".partition("@"))
# ('user', '@', 'domain.com')
# Join
words = ["Hello", "World"]
print(" ".join(words)) # Hello World
print("-".join(["2024", "01", "15"])) # 2024-01-15
print("".join(["a", "b", "c"])) # abc
# String concatenation
s = "Hello" + " " + "World"
parts = ["a"]
parts += "b" # NOT string concat - adds chars to list!Stripping & Padding
strip() removes leading/trailing whitespace by default, or specified characters. zfill() pads with leading zeros (useful for IDs). rjust/ljust/center pad to a specified width with an optional fill character.
# Strip whitespace (or specified chars)
s = " hello "
print(s.strip()) # "hello"
print(s.lstrip()) # "hello "
print(s.rstrip()) # " hello"
# Strip specific characters
print("xxxhelloxxx".strip("x")) # hello
# Padding / centering
print("42".zfill(5)) # 00042
print("hi".rjust(10)) # " hi"
print("hi".ljust(10, "-")) # "hi--------"
print("hi".center(10, "*")) # "****hi****"
# expandtabs
print("a\tb".expandtabs(4)) # "a b"Raw Strings & Escapes
Raw strings (r'...') treat backslashes literally — essential for regex patterns and Windows file paths. Triple-quoted strings preserve newlines. Strings support * operator for repetition and + for concatenation.
# Escape sequences
print("Line1\nLine2") # newline
print("Tab\there") # tab
print("Quote: \"hi\"") # escaped quotes
print("Backslash: \\") # literal backslash
# Raw strings (ignore escapes) - great for regex
path = r"C:\Users\name\file.txt"
regex = r"\d{3}-\d{4}"
print(path) # C:\Users\name\file.txt
# Triple-quoted strings
multi = """
Multiple
lines
"""
# String multiplication
print("ab" * 3) # abababNumbers & Math
Arithmetic Operators
Python has 7 arithmetic operators. / always returns float, // is floor division (rounds toward negative infinity). ** is exponentiation (no ^ — that's XOR). The % operator's result takes the sign of the divisor, unlike C/Java.
# Basic operators
print(7 + 3) # 10 addition
print(7 - 3) # 4 subtraction
print(7 * 3) # 21 multiplication
print(7 / 3) # 2.333... true division (always float)
print(7 // 3) # 2 floor division
print(7 % 3) # 1 modulo (remainder)
print(7 ** 3) # 343 exponentiation
# Floor division with negatives
print(-7 // 3) # -3 (rounds toward negative infinity)
print(-7 % 3) # 2 (result has same sign as divisor)
# Augmented assignment
x = 10
x += 5 # x = x + 5
x **= 2 # x = x ** 2Math Module
The math module provides mathematical functions and constants. All trig functions use radians — convert with math.radians()/degrees(). math.gcd() finds greatest common divisor. For complex numbers, use the cmath module.
import math
# Constants
print(math.pi) # 3.141592653589793
print(math.e) # 2.718281828459045
print(math.inf) # inf
print(math.nan) # nan
# Functions
print(math.sqrt(16)) # 4.0
print(math.pow(2, 10)) # 1024.0
print(math.log(100, 10)) # 2.0 (log base 10)
print(math.log(math.e)) # 1.0 (natural log)
print(math.factorial(5)) # 120
print(math.gcd(12, 8)) # 4
# Rounding
print(math.floor(3.7)) # 3
print(math.ceil(3.2)) # 4
print(math.trunc(-3.7)) # -3 (toward zero)
# Trigonometry (radians)
print(math.sin(math.pi / 2)) # 1.0
print(math.degrees(math.pi)) # 180.0Random Module
The random module uses the Mersenne Twister PRNG — NOT cryptographically secure. Use secrets module for security. random.sample() picks unique items, random.choices() allows duplicates. Set a seed for reproducible results in testing.
import random
# Random integers
print(random.randint(1, 100)) # 1 to 100 inclusive
print(random.randrange(0, 10, 2)) # even number 0,2,4,6,8
# Random floats
print(random.random()) # 0.0 to 1.0
print(random.uniform(1.0, 10.0)) # random float in range
# Choice & sampling
colors = ["red", "green", "blue"]
print(random.choice(colors)) # one random item
print(random.sample(colors, 2)) # 2 unique items
print(random.choices(colors, k=5)) # 5 items (with replacement)
# Shuffle (in-place)
nums = [1, 2, 3, 4, 5]
random.shuffle(nums)
print(nums)
# Reproducible randomness
random.seed(42) # same seed = same sequenceDecimal & Fractions
Use Decimal for financial calculations where float precision errors are unacceptable (e.g., money). Use Fraction for exact rational arithmetic. Both are slower than float but avoid rounding errors. Always construct Decimal from strings, not floats.
from decimal import Decimal, getcontext
from fractions import Fraction
# Float precision issues
print(0.1 + 0.2) # 0.30000000000000004
# Decimal for exact decimal arithmetic
a = Decimal("0.1")
b = Decimal("0.2")
print(a + b) # 0.3 (exact!)
# Set precision
getcontext().prec = 6
print(Decimal(1) / Decimal(7)) # 0.142857
# Fractions for exact rational arithmetic
f1 = Fraction(1, 3)
f2 = Fraction(1, 6)
print(f1 + f2) # 1/2
print(float(f1)) # 0.3333...
# Fraction from string
print(Fraction("3/4")) # 3/4Bitwise Operators
Bitwise operators manipulate individual bits of integers. Python integers have arbitrary precision, so shifts work differently than fixed-width languages. Common uses: flags, masks, low-level protocol parsing. x & (x-1) == 0 checks if x is a power of 2.
# Bitwise operators work on integers
a = 0b1010 # 10
b = 0b1100 # 12
print(a & b) # 8 (0b1000) AND
print(a | b) # 14 (0b1110) OR
print(a ^ b) # 6 (0b0110) XOR
print(~a) # -11 (NOT, two's complement)
print(a << 2) # 40 (left shift, multiply by 4)
print(a >> 1) # 5 (right shift, divide by 2)
# Binary representation
print(bin(10)) # 0b1010
print(hex(255)) # 0xff
print(oct(8)) # 0o10
print(int("1010", 2)) # 10 (parse binary)
# Common tricks
print(5 & 1) # 1 (check odd: nonzero = odd)
print(8 & (8-1)) # 0 (check power of 2)Data Structures
Lists
Lists are Python's most versatile data structure — ordered, mutable, and heterogeneous. append() is O(1), insert(0, x) is O(n). Use collections.deque for fast operations at both ends. sort() is in-place, sorted() returns a new list.
# Lists are ordered, mutable sequences
nums = [1, 2, 3, 4, 5]
mixed = [1, "hello", True, 3.14]
# Adding elements
nums.append(6) # [1,2,3,4,5,6]
nums.insert(0, 0) # [0,1,2,3,4,5,6]
nums.extend([7, 8]) # extend with another list
# Removing elements
nums.remove(0) # remove by value
popped = nums.pop() # remove & return last
popped = nums.pop(0) # remove & return by index
del nums[0] # delete by index
nums.clear() # remove all
# Slicing (same as strings)
nums = [1, 2, 3, 4, 5]
print(nums[1:3]) # [2, 3]
print(nums[::-1]) # [5, 4, 3, 2, 1] reverse
# Sorting
nums.sort() # in-place sort
nums.sort(reverse=True) # descending
sorted_nums = sorted(nums) # returns new listTuples
Tuples are immutable and faster than lists. Use them for fixed collections, multiple return values, and dictionary keys (lists can't be keys). Named tuples provide field names for readability. Single-element tuples need a trailing comma.
# Tuples are ordered, IMMUTABLE sequences
point = (3, 4)
single = (42,) # note the comma for single-element tuple
empty = ()
# Packing & unpacking
coordinates = 10, 20, 30 # packing
x, y, z = coordinates # unpacking
x, y = y, x # swap values!
# Multiple return values
def min_max(nums):
return min(nums), max(nums)
lo, hi = min_max([3, 1, 4, 1, 5])
# Named tuples (readable)
from collections import namedtuple
Point = namedtuple("Point", ["x", "y"])
p = Point(3, 4)
print(p.x, p.y) # 3 4
print(p[0], p[1]) # 3 4
# Tuples are immutable but can contain mutable objects
t = (1, [2, 3])
t[1].append(4) # OK: (1, [2, 3, 4])Dictionaries
Dictionaries are hash maps — O(1) average for lookup/insert/delete. Keys must be hashable (immutable). Since Python 3.7, dicts maintain insertion order. Use get() to avoid KeyError. dict comprehension creates dicts elegantly.
# Dicts are key-value mappings (insertion-ordered since 3.7)
user = {"name": "Alice", "age": 30}
# Access
print(user["name"]) # Alice
print(user.get("email")) # None (no KeyError)
print(user.get("email", "N/A")) # N/A (default)
# Add/update
user["email"] = "[email protected]" # add
user["age"] = 31 # update
user.setdefault("role", "user") # set if missing
# Delete
del user["email"]
val = user.pop("age") # remove & return
# user.clear() # remove all
# Iteration
for key in user: # keys
print(key)
for k, v in user.items(): # key-value pairs
print(k, v)
for v in user.values(): # values
print(v)
# Dict comprehension
squares = {x: x**2 for x in range(5)}
# Merge dicts (3.9+)
merged = {"a": 1} | {"b": 2}Sets
Sets are unordered collections of unique, hashable elements. They excel at membership testing (O(1) vs O(n) for lists) and set algebra (union, intersection, difference). frozenset is immutable and hashable. Order is not guaranteed.
# Sets are unordered collections of unique elements
a = {1, 2, 3, 4}
b = {3, 4, 5, 6}
# Set operations
print(a | b) # union: {1, 2, 3, 4, 5, 6}
print(a & b) # intersection: {3, 4}
print(a - b) # difference: {1, 2}
print(a ^ b) # symmetric difference: {1, 2, 5, 6}
# Methods
a.add(5) # add element
a.discard(10) # remove if present (no error)
a.remove(1) # remove (KeyError if missing)
a.update([6, 7]) # add multiple
# Membership test (O(1) - faster than list)
print(3 in a) # True
# Frozen set (immutable)
fs = frozenset([1, 2, 3])
# Common use: deduplicate
unique = list(set([1, 1, 2, 2, 3])) # [1, 2, 3]Comprehensions
Comprehensions are a Pythonic way to create collections concisely. Generator expressions (parens instead of brackets) are lazy — they produce values on demand, saving memory. Prefer comprehensions over map()/filter() for readability.
# List comprehension
squares = [x**2 for x in range(10)]
evens = [x for x in range(20) if x % 2 == 0]
pairs = [(x, y) for x in range(3) for y in range(3)]
# Dict comprehension
square_map = {x: x**2 for x in range(5)}
# {0: 0, 1: 1, 2: 4, 3: 9, 4: 16}
# Set comprehension
unique_lens = {len(w) for w in ["a", "ab", "abc", "ab"]}
# Generator expression (lazy, memory-efficient)
gen = (x**2 for x in range(1000000))
print(next(gen)) # 0
print(next(gen)) # 1
total = sum(x**2 for x in range(100)) # no extra list
# Nested comprehension (matrix)
matrix = [[i * 3 + j for j in range(3)] for i in range(3)]
# [[0,1,2], [3,4,5], [6,7,8]]Collections Module
The collections module provides specialized containers. Counter counts hashable items. defaultdict auto-creates missing keys. deque offers O(1) append/pop at both ends (vs O(n) for lists). These are essential for clean, efficient code.
from collections import Counter, defaultdict, deque, OrderedDict
# Counter - counting
words = ["apple", "banana", "apple", "cherry", "banana", "apple"]
cnt = Counter(words)
print(cnt) # Counter({'apple': 3, 'banana': 2, 'cherry': 1})
print(cnt.most_common(2)) # [('apple', 3), ('banana', 2)]
# defaultdict - no KeyError
dd = defaultdict(list)
dd["fruits"].append("apple")
dd["fruits"].append("banana")
# dd["vegs"] automatically creates empty list
# deque - fast double-ended queue
dq = deque([1, 2, 3])
dq.appendleft(0) # [0, 1, 2, 3]
dq.append(4) # [0, 1, 2, 3, 4]
dq.popleft() # 0, deque is now [1, 2, 3, 4]
dq.rotate(1) # rotate right
# OrderedDict (less needed since 3.7, dicts are ordered)
od = OrderedDict([("a", 1), ("b", 2)])Control Flow
If / Elif / Else
Python uses if/elif/else — note 'elif' not 'elseif'. Indentation defines blocks. The ternary 'x if cond else y' is an expression. Python treats empty collections, 0, None, and False as falsy — useful for concise conditionals.
score = 85
if score >= 90:
grade = "A"
elif score >= 80:
grade = "B"
elif score >= 70:
grade = "C"
else:
grade = "F"
print(f"Grade: {grade}") # Grade: B
# Conditional expression (ternary)
status = "pass" if score >= 60 else "fail"
# Truthy/falsy values
# Falsy: False, 0, 0.0, "", [], {}, (), None
# Everything else is truthy
if []: # False
print("never")
if [0]: # True (non-empty list)
print("always")For Loops & Iteration
Python's for loop iterates over any iterable. range() generates numbers (exclusive stop). enumerate() pairs items with indices. zip() iterates multiple sequences in parallel. Use .items() to iterate dict key-value pairs.
# range(start, stop, step)
for i in range(5): # 0, 1, 2, 3, 4
print(i)
for i in range(2, 10, 2): # 2, 4, 6, 8
print(i)
for i in range(10, 0, -1): # countdown
print(i)
# Iterate over collections
fruits = ["apple", "banana", "cherry"]
for fruit in fruits:
print(fruit)
# enumerate for index + value
for idx, fruit in enumerate(fruits):
print(f"{idx}: {fruit}")
# zip to iterate multiple sequences
names = ["Alice", "Bob"]
ages = [30, 25]
for name, age in zip(names, ages):
print(f"{name}: {age}")
# Iterate dict
user = {"name": "Alice", "age": 30}
for key, value in user.items():
print(f"{key} = {value}")While Loops & Break/Continue
while loops repeat while a condition is true. break exits the loop immediately, continue skips to the next iteration. The for/else construct runs the else block only if the loop wasn't broken. pass is a no-op placeholder for empty blocks.
# Basic while
count = 0
while count < 5:
print(count)
count += 1
# break - exit loop
while True:
cmd = input("> ")
if cmd == "quit":
break
print(f"You said: {cmd}")
# continue - skip to next iteration
for i in range(10):
if i % 2 == 0:
continue # skip even numbers
print(i) # prints 1, 3, 5, 7, 9
# else clause (runs if no break)
for i in range(5):
if i == 10:
break
else:
print("Loop completed without break")
# pass - do nothing (placeholder)
for i in range(5):
pass # TODO: implementMatch Statement (Python 3.10+)
The match statement (Python 3.10+) is powerful structural pattern matching, far beyond C's switch. It can match sequences, mappings, class instances, and bind variables. The _ pattern is a wildcard (default). Guards with 'if' add conditions.
# Structural pattern matching (like switch)
def handle_command(cmd):
match cmd.split():
case ["quit"]:
return "Goodbye"
case ["hello", name]:
return f"Hello, {name}!"
case ["move", direction] if direction in "NSEW":
return f"Moving {direction}"
case ["add", x, y]:
return int(x) + int(y)
case _:
return "Unknown command"
print(handle_command("hello Alice")) # Hello, Alice!
# Matching data structures
match point:
case (0, 0):
print("origin")
case (0, y):
print(f"on y-axis at {y}")
case (x, 0):
print(f"on x-axis at {x}")
case (x, y):
print(f"at ({x}, {y})")Iterators & Generators
Generators produce values lazily using yield — they don't compute all values upfront, saving memory. They implement the iterator protocol (iter() and next()). Once exhausted, they're done. Use generators for large/infinite sequences, pipelines, and streaming data.
# Iterator protocol
nums = [1, 2, 3]
it = iter(nums)
print(next(it)) # 1
print(next(it)) # 2
print(next(it)) # 3
# next(it) # StopIteration
# Generator function (uses yield)
def count_up_to(n):
count = 1
while count <= n:
yield count
count += 1
for num in count_up_to(5):
print(num) # 1, 2, 3, 4, 5
# Infinite generator
def fibonacci():
a, b = 0, 1
while True:
yield a
a, b = b, a + b
fib = fibonacci()
print(next(fib)) # 0
print(next(fib)) # 1
print(next(fib)) # 1
print(next(fib)) # 2
# Generator expression
squares = (x**2 for x in range(10))Functions
Defining & Calling Functions
Functions are defined with def. Default arguments use =. Python supports keyword arguments for clarity. Functions can return multiple values (as a tuple). Docstrings (triple-quoted) document functions and are accessible via help().
# Basic function
def greet(name):
return f"Hello, {name}!"
print(greet("Alice")) # Hello, Alice!
# Default arguments
def greet(name, greeting="Hello"):
return f"{greeting}, {name}!"
print(greet("Bob")) # Hello, Bob!
print(greet("Bob", "Hi")) # Hi, Bob!
# Keyword arguments
print(greet(name="Carol", greeting="Hey"))
# Return multiple values (tuple)
def stats(nums):
return min(nums), max(nums), sum(nums) / len(nums)
lo, hi, avg = stats([1, 2, 3, 4, 5])
# Docstrings
def add(a, b):
"""Add two numbers and return the result.
Args:
a: First number
b: Second number
Returns:
Sum of a and b
"""
return a + bArguments: *args & **kwargs
*args collects extra positional arguments into a tuple, **kwargs collects extra keyword arguments into a dict. The names args/kwargs are convention. You can unpack sequences with * and dicts with ** when calling functions. Order: positional, *args, keyword, **kwargs.
# *args - variable positional arguments (tuple)
def sum_all(*args):
return sum(args)
print(sum_all(1, 2, 3)) # 6
print(sum_all(1, 2, 3, 4, 5)) # 15
# **kwargs - variable keyword arguments (dict)
def print_info(**kwargs):
for key, value in kwargs.items():
print(f"{key}: {value}")
print_info(name="Alice", age=30, role="admin")
# Combining all
def func(a, b, *args, **kwargs):
print(f"a={a}, b={b}")
print(f"args={args}")
print(f"kwargs={kwargs}")
func(1, 2, 3, 4, x=5, y=6)
# a=1, b=2, args=(3, 4), kwargs={'x': 5, 'y': 6}
# Unpacking arguments
nums = [1, 2, 3]
print(sum_all(*nums)) # unpack list as args
opts = {"name": "Alice", "age": 30}
print_info(**opts) # unpack dict as kwargsLambda & Higher-Order Functions
Lambdas are limited to a single expression — use def for complex logic. They shine as arguments to higher-order functions like sorted(), map(), filter(). However, list comprehensions are often more readable than map/filter. reduce() lives in functools.
# Lambda - anonymous function (single expression)
square = lambda x: x ** 2
print(square(5)) # 25
# Common with sorted, map, filter, reduce
students = [("Alice", 85), ("Bob", 92), ("Carol", 78)]
# Sort by score (key function)
sorted_by_score = sorted(students, key=lambda s: s[1])
# [('Carol', 78), ('Alice', 85), ('Bob', 92)]
# map - apply function to each item
nums = [1, 2, 3, 4, 5]
doubled = list(map(lambda x: x * 2, nums))
# [2, 4, 6, 8, 10]
# filter - keep items where function returns True
evens = list(filter(lambda x: x % 2 == 0, nums))
# [2, 4]
# reduce - accumulate to single value
from functools import reduce
product = reduce(lambda a, b: a * b, nums)
# 120 (1*2*3*4*5)
# Prefer comprehensions over map/filter
doubled = [x * 2 for x in nums] # more Pythonic
evens = [x for x in nums if x % 2 == 0]Decorators
Decorators wrap functions to add behavior without modifying the original code. The @syntax is syntactic sugar. Decorators with arguments need an extra nesting level. Always use functools.wraps to preserve the original function's metadata (name, docstring).
# A decorator modifies a function's behavior
def uppercase_result(func):
def wrapper(*args, **kwargs):
result = func(*args, **kwargs)
return result.upper()
return wrapper
@uppercase_result
def greet(name):
return f"hello, {name}"
print(greet("alice")) # HELLO, ALICE
# Decorator with arguments
def repeat(times):
def decorator(func):
def wrapper(*args, **kwargs):
for _ in range(times):
result = func(*args, **kwargs)
return result
return wrapper
return decorator
@repeat(3)
def say_hi():
print("Hi!")
say_hi() # prints "Hi!" three times
# Practical: timing decorator
import time
def timer(func):
def wrapper(*args, **kwargs):
start = time.time()
result = func(*args, **kwargs)
print(f"{func.__name__} took {time.time() - start:.4f}s")
return result
return wrapper
# Use functools.wraps to preserve metadata
from functools import wraps
def my_decorator(func):
@wraps(func)
def wrapper(*args, **kwargs):
return func(*args, **kwargs)
return wrapperScope & Closures
Python resolves names using LEGB scope order: Local, Enclosing, Global, Built-in. Use 'global' to rebind a global variable inside a function. Use 'nonlocal' (Python 3) to modify a variable in an enclosing scope. Closures remember their enclosing scope.
# LEGB rule: Local, Enclosing, Global, Built-in
x = "global"
def outer():
x = "enclosing"
def inner():
x = "local"
print(x) # local
inner()
print(x) # enclosing
outer()
print(x) # global
# global keyword - modify global variable
count = 0
def increment():
global count
count += 1
# nonlocal keyword - modify enclosing variable
def make_counter():
count = 0
def counter():
nonlocal count
count += 1
return count
return counter
c = make_counter()
print(c()) # 1
print(c()) # 2
print(c()) # 3OOP & Classes
Classes & Objects
Classes bundle data (attributes) and behavior (methods). __init__ is the constructor. self refers to the instance (like 'this' in other languages). Class variables are shared; instance variables are per-object. __str__ is for users, __repr__ is for developers.
class Dog:
# Class variable (shared by all instances)
species = "Canis familiaris"
# Constructor
def __init__(self, name, age):
# Instance variables
self.name = name
self.age = age
# Instance method
def bark(self):
return f"{self.name} says Woof!"
# String representation
def __str__(self):
return f"Dog({self.name}, {self.age})"
# Official representation (for debugging)
def __repr__(self):
return f"Dog(name='{self.name}', age={self.age})"
# Create instances
buddy = Dog("Buddy", 3)
lucy = Dog("Lucy", 5)
print(buddy.bark()) # Buddy says Woof!
print(buddy.name) # Buddy
print(buddy.species) # Canis familiaris
print(str(buddy)) # Dog(Buddy, 3)Inheritance & Polymorphism
Inheritance lets classes reuse and extend behavior. Python supports multiple inheritance with MRO (Method Resolution Order) to resolve conflicts. Polymorphism allows treating different types uniformly. Use isinstance() for type checking, not type().
class Animal:
def __init__(self, name):
self.name = name
def speak(self):
raise NotImplementedError("Subclass must implement")
class Dog(Animal):
def speak(self):
return f"{self.name} says Woof!"
class Cat(Animal):
def speak(self):
return f"{self.name} says Meow!"
# Polymorphism - same interface, different behavior
def animal_sound(animal):
print(animal.speak())
animals = [Dog("Buddy"), Cat("Whiskers")]
for a in animals:
animal_sound(a)
# Buddy says Woof!
# Whiskers says Meow!
# Multiple inheritance
class Swimmer:
def swim(self):
return "swimming"
class Flyer:
def fly(self):
return "flying"
class Duck(Animal, Swimmer, Flyer):
pass
duck = Duck("Donald")
print(duck.swim()) # swimming
print(duck.fly()) # flying
# Check inheritance
print(isinstance(duck, Animal)) # True
print(issubclass(Dog, Animal)) # TrueProperties & Encapsulation
Python has no true private/protected — it uses conventions. A single underscore _ means 'internal'. Double underscore __ triggers name mangling (not true privacy). @property turns methods into attributes with getters/setters, enabling validation and computed properties.
class Temperature:
def __init__(self, celsius=0):
self.celsius = celsius # uses setter below
# Getter
@property
def celsius(self):
return self._celsius
# Setter
@celsius.setter
def celsius(self, value):
if value < -273.15:
raise ValueError("Below absolute zero!")
self._celsius = value
# Computed property
@property
def fahrenheit(self):
return self._celsius * 9/5 + 32
@fahrenheit.setter
def fahrenheit(self, value):
self.celsius = (value - 32) * 5/9
temp = Temperature(25)
print(temp.fahrenheit) # 77.0
temp.fahrenheit = 100
print(temp.celsius) # 37.78...
# Name conventions:
# _name - protected (convention, not enforced)
# __name - private (name mangling: _ClassName__name)
# __name__ - dunder (reserved by Python)Class & Static Methods
@staticmethod is just a function in a class namespace — no implicit first argument. @classmethod receives the class (cls) as first argument, useful for alternative constructors (factory methods) and inheritance-aware behavior. Use classmethod for constructors, staticmethod for utility functions.
class MathUtils:
pi = 3.14159
# Static method - no self/cls, lives in class namespace
@staticmethod
def add(a, b):
return a + b
# Class method - receives the class as first argument
@classmethod
def circle_area(cls, radius):
return cls.pi * radius ** 2
# Alternative constructor (common classmethod use)
@classmethod
def from_diameter(cls, diameter):
return cls() # would configure instance
# Static: called on class or instance, no special first arg
print(MathUtils.add(2, 3)) # 5
# Class: often used for alternative constructors
print(MathUtils.circle_area(5)) # 78.54...
# Factory pattern with classmethod
class Point:
def __init__(self, x, y):
self.x = x
self.y = y
@classmethod
def origin(cls):
return cls(0, 0)
@classmethod
def from_tuple(cls, coords):
return cls(*coords)
p1 = Point.origin()
p2 = Point.from_tuple((3, 4))Magic Methods (Dunder)
Magic methods (dunder methods) implement operator overloading and protocol behavior. __add__ for +, __eq__ for ==, __len__ for len(), __iter__ for iteration/unpacking. They let your objects work with Python's built-in syntax and functions naturally.
class Vector:
def __init__(self, x, y):
self.x = x
self.y = y
# String representation
def __str__(self):
return f"Vector({self.x}, {self.y})"
# Operator overloading
def __add__(self, other):
return Vector(self.x + other.x, self.y + other.y)
def __mul__(self, scalar):
return Vector(self.x * scalar, self.y * scalar)
# Equality
def __eq__(self, other):
return self.x == other.x and self.y == other.y
# Length
def __len__(self):
return int((self.x**2 + self.y**2) ** 0.5)
# Make it iterable
def __iter__(self):
yield self.x
yield self.y
# Index access
def __getitem__(self, index):
return (self.x, self.y)[index]
v1 = Vector(2, 3)
v2 = Vector(4, 5)
print(v1 + v2) # Vector(6, 8)
print(v1 * 3) # Vector(6, 9)
print(v1 == Vector(2, 3)) # True
print(len(v1)) # 3
x, y = v1 # unpacking via __iter__Error Handling
Try / Except / Finally
try/except/else/finally: try runs risky code, except catches errors, else runs if no exception, finally always runs (cleanup). Catch specific exceptions, not bare 'except:'. The else block is useful when cleanup should only happen on success. Exception is the base for most catchable errors.
# Basic exception handling
try:
result = 10 / 0
except ZeroDivisionError as e:
print(f"Error: {e}") # division by zero
finally:
print("This always runs")
# Multiple exception types
try:
value = int("abc")
except (ValueError, TypeError) as e:
print(f"Conversion error: {e}")
# Different handlers for different exceptions
try:
f = open("nonexistent.txt")
data = f.read()
except FileNotFoundError:
print("File not found")
except PermissionError:
print("No permission")
except Exception as e:
print(f"Unexpected: {e}")
else:
print("No exception occurred")
f.close()
finally:
print("Cleanup (always runs)")
# Exception hierarchy
# BaseException
# ├── SystemExit
# ├── KeyboardInterrupt
# └── Exception
# ├── ValueError
# ├── TypeError
# ├── KeyError
# └── ...Raising Exceptions
Use raise to throw exceptions. 'raise' alone re-raises the current exception (in an except block). 'raise X from Y' chains exceptions, preserving the original cause. Always raise specific exception types. Avoid using exceptions for normal control flow.
# Raise an exception
def divide(a, b):
if b == 0:
raise ZeroDivisionError("Cannot divide by zero!")
return a / b
# Re-raise the current exception
def process(data):
try:
return parse(data)
except ValueError:
print("Logging parse error...")
raise # re-raises the same exception
# Raise with context (from)
try:
int("abc")
except ValueError as e:
raise RuntimeError("Failed to process input") from e
# Common built-in exceptions
raise ValueError("invalid value")
raise TypeError("wrong type")
raise KeyError("missing key")
raise IndexError("out of range")
raise RuntimeError("something went wrong")
raise NotImplementedError("override this")
raise FileNotFoundError("no such file")
# Exception with custom args
class ValidationError(Exception):
pass
raise ValidationError("field is required", "email")Custom Exceptions
Create custom exceptions by subclassing Exception (or a more specific built-in). Design a hierarchy so callers can catch at the right level. Add custom attributes to carry context. Inherit from Exception, not BaseException (which includes SystemExit/KeyboardInterrupt).
# Custom exception hierarchy
class AppError(Exception):
"""Base exception for the application."""
pass
class DatabaseError(AppError):
def __init__(self, message, query=None):
super().__init__(message)
self.query = query
class ValidationError(AppError):
def __init__(self, field, message):
super().__init__(f"{field}: {message}")
self.field = field
self.message = message
class AuthenticationError(AppError):
pass
# Usage
def login(username, password):
if not username:
raise ValidationError("username", "is required")
if password != "secret":
raise AuthenticationError("Invalid credentials")
# Catching by hierarchy
try:
login("", "x")
except ValidationError as e:
print(f"Validation failed: {e.field}")
except AppError as e:
print(f"App error: {e}")
# Access exception info
import traceback
try:
1 / 0
except:
traceback.print_exc()
print(repr(sys.exc_info()[1]))Context Managers (with statement)
Context managers (the 'with' statement) guarantee cleanup via __enter__ and __exit__. They're essential for resources like files, locks, and database connections. contextlib.contextmanager simplifies creation using a generator. __exit__ can suppress exceptions by returning True.
# Context managers handle setup and cleanup
with open("file.txt") as f:
content = f.read()
# file is automatically closed, even if an error occurs
# Multiple context managers
with open("input.txt") as fin, open("output.txt", "w") as fout:
fout.write(fin.read())
# Creating a context manager (class-based)
class Timer:
def __enter__(self):
import time
self.start = time.time()
return self
def __exit__(self, exc_type, exc_val, exc_tb):
import time
self.elapsed = time.time() - self.start
print(f"Elapsed: {self.elapsed:.4f}s")
return False # don't suppress exceptions
with Timer() as t:
# code to time
sum(range(1000000))
# contextlib for simpler context managers
from contextlib import contextmanager
@contextmanager
def open_db(url):
db = connect(url)
try:
yield db
finally:
db.close()
with open_db("localhost") as db:
db.query("SELECT 1")Assertions & Logging
assert statements are for debugging invariants — they're stripped when Python runs with -O (optimize). Never use assertions for input validation. Use the logging module instead of print() for production code — it supports levels, formatting, and output destinations.
# Assertions - for debugging (removed with -O flag)
def divide(a, b):
assert b != 0, "Divisor cannot be zero"
return a / b
# Never use assertions for data validation (they can be disabled)
# Use them for invariant checks during development
# Logging (better than print for production)
import logging
# Configure logging
logging.basicConfig(
level=logging.DEBUG,
format="%(asctime)s [%(levelname)s] %(message)s"
)
logger = logging.getLogger(__name__)
logger.debug("Detailed info for debugging")
logger.info("Confirmation things are working")
logger.warning("Something unexpected happened")
logger.error("A serious problem")
logger.critical("A fatal error")
# Log exceptions with traceback
try:
1 / 0
except:
logger.exception("Failed to divide") # includes tracebackFile I/O
Reading Files
Always use 'with' to open files — it automatically closes them even if an error occurs. Specify encoding='utf-8' to avoid platform-dependent encoding issues. For large files, iterate line-by-line instead of read() to save memory. readlines() loads the entire file into memory.
# Read entire file
with open("file.txt", "r", encoding="utf-8") as f:
content = f.read()
print(content)
# Read line by line (memory-efficient for large files)
with open("file.txt", "r") as f:
for line in f:
print(line.strip()) # strip removes trailing newline
# Read all lines into a list
with open("file.txt") as f:
lines = f.readlines() # ['line1\n', 'line2\n', ...]
# Read specific number of characters
with open("file.txt") as f:
chunk = f.read(100) # first 100 chars
# File modes:
# "r" read (default)
# "w" write (truncate)
# "a" append
# "x" exclusive create (fails if exists)
# "b" binary mode (e.g., "rb", "wb")
# "+" read and write (e.g., "r+")Writing Files
Mode 'w' truncates the file (deletes content); use 'a' to append. writelines() doesn't add newlines — add them manually. Use 'rb'/'wb' for binary files (images, etc.). seek() moves the cursor; tell() returns its position. Always specify encoding for text files.
# Write text (overwrites existing)
with open("output.txt", "w") as f:
f.write("First line\n")
f.write("Second line\n")
# writelines doesn't add newlines
f.writelines(["line3\n", "line4\n"])
# Append to a file
with open("log.txt", "a") as f:
f.write("New log entry\n")
# Write binary data
with open("data.bin", "wb") as f:
f.write(b"\x00\x01\x02\x03")
# Read and write simultaneously
with open("file.txt", "r+") as f:
content = f.read()
f.seek(0) # move to beginning
f.write("Updated") # overwrite
f.truncate() # cut off remaining
# Check if file exists
import os
if os.path.exists("file.txt"):
print("File exists")Path Handling (pathlib)
pathlib (Python 3.4+) is the modern, object-oriented way to handle paths — prefer it over os.path. The / operator joins paths platform-independently. Path objects have read_text()/write_text() methods that handle open/close for you. rglob() recursively searches.
from pathlib import Path
# Create Path objects (preferred over os.path)
p = Path("src/main.py")
home = Path.home() # /home/user or C:\Users\user
cwd = Path.cwd() # current working directory
# Path components
print(p.name) # main.py
print(p.stem) # main
print(p.suffix) # .py
print(p.parent) # src
print(p.parts) # ('src', 'main.py')
# Joining paths (use / operator)
config = home / ".config" / "app" / "config.json"
# Existence and type
print(p.exists()) # True/False
print(p.is_file())
print(p.is_dir())
# Listing directories
for f in Path(".").iterdir():
print(f)
# Glob patterns
for py_file in Path(".").rglob("*.py"):
print(py_file)
# Create directories
Path("new/dir").mkdir(parents=True, exist_ok=True)
# Read/write (Path methods)
content = Path("file.txt").read_text()
Path("output.txt").write_text("Hello!")JSON
json.dumps() serializes to string, json.loads() deserializes. dump()/load() work with files. Use indent for pretty-printing. Custom objects need a default serializer. JSON only supports basic types — use default= for datetime and other complex objects.
import json
# Python dict to JSON string
data = {"name": "Alice", "age": 30, "skills": ["Python", "SQL"]}
json_str = json.dumps(data, indent=2)
print(json_str)
# JSON string to Python dict
parsed = json.loads('{"name": "Bob", "active": true}')
print(parsed["name"]) # Bob
print(parsed["active"]) # True (Python bool)
# Write JSON to file
with open("data.json", "w") as f:
json.dump(data, f, indent=2, ensure_ascii=False)
# Read JSON from file
with open("data.json") as f:
loaded = json.load(f)
# Custom serialization (e.g., datetime)
from datetime import datetime
def json_default(obj):
if isinstance(obj, datetime):
return obj.isoformat()
raise TypeError
json.dumps({"time": datetime.now()}, default=json_default)
# Type mapping:
# JSON object <-> Python dict
# JSON array <-> Python list
# JSON string <-> Python str
# JSON number <-> Python int/float
# JSON boolean <-> Python bool
# JSON null <-> Python NoneCSV & Other Formats
The csv module handles CSV with proper quoting. Use newline='' when opening CSV files on Windows. DictReader/DictWriter work with column names. pickle can serialize any Python object but is Python-specific and insecure — never unpickle data from untrusted sources.
import csv
# Write CSV
with open("data.csv", "w", newline="") as f:
writer = csv.writer(f)
writer.writerow(["name", "age", "city"])
writer.writerows([
["Alice", 30, "NYC"],
["Bob", 25, "LA"],
])
# Read CSV
with open("data.csv") as f:
reader = csv.reader(f)
header = next(reader) # first row
for row in reader:
print(row) # ['Alice', '30', 'NYC']
# DictReader/DictWriter (column access by name)
with open("data.csv") as f:
reader = csv.DictReader(f)
for row in reader:
print(row["name"], row["age"])
# Pickle (Python-specific, can store any object)
import pickle
with open("data.pkl", "wb") as f:
pickle.dump({"complex": [1, 2, {"a": 3}]}, f)
with open("data.pkl", "rb") as f:
obj = pickle.load(f)
# WARNING: pickle is insecure - never unpickle untrusted data!Modules & Packages
Importing Modules
Imports bring in modules. 'import X' keeps the namespace clean. 'from X import Y' is convenient but can cause name collisions. Aliases (import X as Y) are common for libraries with conventions (np, pd). Avoid 'from X import *' — it pollutes the namespace.
# Import entire module
import math
print(math.sqrt(16))
# Import specific names
from datetime import datetime, timedelta
now = datetime.now()
# Import with alias
import numpy as np
import pandas as pd
# Import all names (discouraged - pollutes namespace)
# from os import *
# Conditional import (try/except)
try:
import cjson as json
except ImportError:
import json
# Check what's in a module
import os
print(dir(os)) # list all attributes
print(os.__file__) # module location
print(os.__name__) # module name
# Reload a module (during development)
import importlib
importlib.reload(my_module)Creating Modules & Packages
A module is a .py file; a package is a directory with __init__.py. The __init__.py file can be empty or set up the package. __all__ in __init__.py controls what 'from package import *' exports. Modern Python (3.3+) supports namespace packages without __init__.py.
# A module is just a .py file
# mymath.py
def add(a, b):
return a + b
PI = 3.14159
# A package is a directory with __init__.py
# mypackage/
# __init__.py
# module1.py
# module2.py
# subpackage/
# __init__.py
# module3.py
# __init__.py can be empty or contain package initialization
# mypackage/__init__.py
from .module1 import ClassA
from .module2 import func_b
__version__ = "1.0.0"
__all__ = ["ClassA", "func_b"]
# Using the package
from mypackage import ClassA
from mypackage.subpackage import module3
# __all__ controls 'from package import *'
# Without __all__, * imports only what's in __init__.py__name__ == '__main__'
The if __name__ == '__main__' idiom lets a file serve as both a script and a module. When run directly, __name__ is '__main__'; when imported, it's the module name. This pattern is essential for making reusable modules that can also be executed standalone.
# script.py
def main():
print("Running main")
def helper():
print("Helper function")
if __name__ == "__main__":
# This code only runs when the file is executed directly
# NOT when imported as a module
main()
# When you run: python script.py
# __name__ is "__main__" -> main() runs
# When you: import script
# __name__ is "script" -> main() does NOT run
# This lets the module be both a script and an importable library
# Common pattern for CLI tools
def main():
import argparse
parser = argparse.ArgumentParser()
parser.add_argument("--name", required=True)
args = parser.parse_args()
print(f"Hello, {args.name}!")
if __name__ == "__main__":
main()Standard Library Highlights
Python's standard library is huge and batteries-included. os/sys for system interaction, datetime for dates, collections for specialized containers, itertools for iterator tools, functools for functional programming. Explore the docs at docs.python.org/3/library/.
# os - operating system interface
import os
os.getcwd() # current directory
os.listdir(".") # list files
os.environ.get("HOME") # environment variables
# sys - system-specific
import sys
sys.argv # command-line arguments
sys.exit(0) # exit with status code
sys.path # module search path
# datetime - date and time
from datetime import datetime, timedelta
now = datetime.now()
# collections - specialized containers
from collections import Counter, defaultdict, deque
# itertools - iterator tools
from itertools import chain, cycle, repeat, product
# functools - higher-order functions
from functools import lru_cache, reduce, partial
# typing - type hints
from typing import List, Dict, Optional, Union, Any
# pathlib - path handling
from pathlib import Path
# subprocess - run external commands
import subprocess
result = subprocess.run(["ls", "-l"], capture_output=True, text=True)Pip & Virtual Environments
Always use virtual environments to isolate project dependencies. venv is built-in; alternatives include virtualenv, conda, and uv. Pin versions in requirements.txt for reproducibility. Never install packages globally with --user or as root — use a venv.
# Create a virtual environment
# $ python -m venv venv
# Activate it
# Windows: venv\Scripts\activate
# Unix: source venv/bin/activate
# Install packages
# $ pip install requests
# $ pip install requests==2.28.0
# $ pip install "requests>=2.25,<3.0"
# Install from requirements file
# $ pip install -r requirements.txt
# requirements.txt example:
# requests==2.31.0
# numpy>=1.21.0
# pandas~=2.0.0 # compatible release
# List installed packages
# $ pip list
# $ pip freeze > requirements.txt
# Uninstall
# $ pip uninstall requests
# Show package info
# $ pip show requests
# Modern alternative: uv (faster)
# $ uv pip install requestsDate & Time
datetime Module
The datetime module provides date, time, datetime, and timedelta classes. datetime.now() returns local time; use datetime.now(timezone.utc) for UTC. weekday() returns 0-6 (Mon-Sun). Always use timezone-aware datetimes in production to avoid ambiguity.
from datetime import datetime, date, time, timedelta
# Current date and time
now = datetime.now() # local time
utc_now = datetime.utcnow() # UTC (deprecated in 3.12)
utc = datetime.now(timezone.utc) # preferred
# Current date
today = date.today()
# Create specific date/time
dt = datetime(2024, 1, 15, 10, 30, 0)
d = date(2024, 1, 15)
t = time(10, 30, 0)
# Access components
print(now.year, now.month, now.day)
print(now.hour, now.minute, now.second)
print(now.weekday()) # 0=Monday, 6=Sunday
# From timestamp
ts = 1705315200
dt = datetime.fromtimestamp(ts)
# To timestamp
print(datetime.now().timestamp())Formatting & Parsing
strftime (string format time) converts datetime to string; strptime (string parse time) converts string to datetime. ISO 8601 format (isoformat/fromisoformat) is the best choice for storing dates — it's unambiguous and sortable. Memorize common codes: %Y %m %d %H %M %S.
from datetime import datetime
# Format datetime to string (strftime)
dt = datetime(2024, 1, 15, 10, 30)
print(dt.strftime("%Y-%m-%d")) # 2024-01-15
print(dt.strftime("%Y/%m/%d %H:%M")) # 2024/01/15 10:30
print(dt.strftime("%B %d, %Y")) # January 15, 2024
print(dt.strftime("%A")) # Monday
# Parse string to datetime (strptime)
dt = datetime.strptime("2024-01-15", "%Y-%m-%d")
dt = datetime.strptime("15/01/2024 10:30", "%d/%m/%Y %H:%M")
# Common format codes:
# %Y year (2024) %m month (01)
# %d day (15) %H hour (14)
# %M minute (30) %S second (00)
# %B month name %b month abbrev
# %A weekday name %a weekday abbrev
# %I 12-hour %p AM/PM
# %j day of year %U week number
# ISO format (recommended for storage)
iso = dt.isoformat() # "2024-01-15T10:30:00"
dt = datetime.fromisoformat("2024-01-15T10:30:00")timedelta & Arithmetic
timedelta represents a duration. You can add/subtract timedeltas from datetimes and subtract two datetimes to get a timedelta. timedelta normalizes: days=1, hours=25 becomes days=2, hours=1. total_seconds() gives the entire duration in seconds.
from datetime import datetime, timedelta
now = datetime.now()
# Add/subtract time
tomorrow = now + timedelta(days=1)
last_week = now - timedelta(weeks=1)
in_2_hours = now + timedelta(hours=2)
in_90_days = now + timedelta(days=90)
# Difference between dates
date1 = datetime(2024, 1, 15)
date2 = datetime(2024, 6, 18)
diff = date2 - date1
print(diff.days) # 155
print(diff.total_seconds())
# timedelta components
td = timedelta(days=5, hours=3, minutes=30)
print(td.days) # 5
print(td.seconds) # 12600 (3h 30m in seconds)
print(td.total_seconds())
# Comparisons
if now > date1:
print("now is later")
# Business day calculation (using numpy)
# import numpy as np
# business_days = np.busday_count(date1.date(), date2.date())Timezones
Always use timezone-aware datetimes (Python 3.9+ ZoneInfo is preferred over pytz). Store dates in UTC and convert to local time only for display. Naive datetimes (without tzinfo) cause subtle bugs. ZoneInfo uses the IANA timezone database, handling DST automatically.
from datetime import datetime, timezone, timedelta
# Timezone-aware datetime
utc_time = datetime.now(timezone.utc)
print(utc_time) # 2024-01-15 10:30:00+00:00
# Create a timezone (offset-based)
tz_ny = timezone(timedelta(hours=-5), "EST")
ny_time = datetime.now(tz_ny)
# Convert between timezones
utc_time = datetime.now(timezone.utc)
ny_time = utc_time.astimezone(timezone(timedelta(hours=-5)))
tokyo_time = utc_time.astimezone(timezone(timedelta(hours=9)))
# Use zoneinfo (Python 3.9+) for IANA timezones
from zoneinfo import ZoneInfo
tz = ZoneInfo("America/New_York")
dt = datetime.now(tz)
print(dt.tzname()) # EST or EDT
# Common IANA timezones:
# "UTC"
# "America/New_York", "America/Los_Angeles"
# "Europe/London", "Europe/Paris"
# "Asia/Tokyo", "Asia/Shanghai"
# "Australia/Sydney"
# Best practice: store UTC, convert for display
utc_stored = datetime.now(timezone.utc)
local_display = utc_stored.astimezone(ZoneInfo("Asia/Shanghai"))Regular Expressions
re Module Basics
re.search() finds the first match anywhere; re.match() only at the start; re.fullmatch() requires the entire string to match. Use raw strings (r'...') for patterns to avoid backslash escaping issues. Match objects provide group(), start(), end(), and span().
import re
# re.search - find first match anywhere in string
m = re.search(r"\d{4}", "Order #2024 was placed")
if m:
print(m.group()) # 2024
print(m.start(), m.end()) # 9 13
# re.match - match at beginning of string
m = re.match(r"Hello", "Hello, World")
print(m.group()) # Hello
# re.fullmatch - entire string must match
m = re.fullmatch(r"\d+", "12345")
print(bool(m)) # True
# re.findall - all matches as list
emails = re.findall(r"\S+@\S+", text)
numbers = re.findall(r"\d+", "a1b22c333") # ['1', '22', '333']
# re.finditer - all matches as iterator (with positions)
for m in re.finditer(r"\w+", "Hello World"):
print(m.group(), m.span())
# Match object methods
m = re.search(r"(\w+)@(\w+)", "[email protected]")
print(m.group()) # user@domain (whole match)
print(m.group(1)) # user (first group)
print(m.group(2)) # domain (second group)
print(m.groups()) # ('user', 'domain')Pattern Syntax
Regex syntax: [] for character classes, \d \w \s for common sets, quantifiers (* + ? {}) for repetition, ^ $ \b for anchors, () for groups, | for alternation. Use raw strings (r'...') so backslashes are literal. Greedy quantifiers match as much as possible; add ? for lazy (e.g., *?).
import re
# Character classes
re.findall(r"[aeiou]", "hello") # vowels
re.findall(r"[^aeiou]", "hello") # non-vowels
re.findall(r"[a-z]", "Hello123") # lowercase
re.findall(r"[A-Za-z0-9]", "Hi-1!") # alphanumeric
# Predefined classes
# . any char except newline
# \d digit [0-9] \D non-digit
# \w word char [a-zA-Z0-9_] \W non-word
# \s whitespace \S non-whitespace
# Quantifiers
# * 0 or more
# + 1 or more
# ? 0 or 1
# {n} exactly n
# {n,} n or more
# {n,m} between n and m
re.findall(r"\d{3}", "1234567") # ['123', '456']
re.findall(r"\d{2,4}", "12345678") # ['1234', '5678']
# Anchors
# ^ start of string $ end of string
# \b word boundary
re.findall(r"^\w+", "Hello World") # ['Hello']
re.findall(r"\b\w+\b", "hi there") # ['hi', 'there']
# Groups & alternation
re.findall(r"(cat|dog)", "cat and dog") # ['cat', 'dog']
re.findall(r"(\w+)@(\w+\.\w+)", "[email protected]")Substitution & Splitting
re.sub() replaces matches — use backreferences (\1, \2) to reference groups, or a function for dynamic replacement. re.split() is more powerful than str.split() — it accepts regex patterns. Capture groups in the pattern are included in the result. Use re.subn() to get (result, count).
import re
# re.sub - replace matches
result = re.sub(r"\d+", "#", "a1b22c333")
# 'a#b#c#'
# Replace with count limit
result = re.sub(r"\d", "X", "a1b2c3", count=2)
# 'aXbXc3'
# Use backreferences in replacement
result = re.sub(r"(\w+)@(\w+)", r"\2.\1", "user@domain")
# 'domain.user'
# Use function as replacement
def upper(m):
return m.group().upper()
result = re.sub(r"\b[a-z]", upper, "hello world")
# 'Hello World' (capitalize first letter of each word)
# re.split - split by pattern
parts = re.split(r"[,;\s]+", "a, b; c d")
# ['a', 'b', 'c', 'd']
# Split with capture groups (keeps delimiters)
parts = re.split(r"([,;])", "a,b;c")
# ['a', ',', 'b', ';', 'c']
# Split with maxsplit
parts = re.split(r",", "a,b,c,d", maxsplit=2)
# ['a', 'b', 'c,d']Compilation & Flags
Compile patterns with re.compile() when using them repeatedly — it's faster. Flags modify behavior: IGNORECASE, MULTILINE, DOTALL, VERBOSE (allows comments/whitespace in patterns). Named groups (?P<name>...) improve readability. Lookahead (?=) and lookbehind (?<=) match without consuming.
import re
# Compile pattern for reuse (faster when used many times)
email_re = re.compile(r"^[\w.+-]+@([\w-]+\.)+[\w-]+$")
print(email_re.match("[email protected]")) # match object
print(email_re.match("invalid")) # None
# Common flags
re.IGNORECASE # case-insensitive
re.MULTILINE # ^ and $ match line boundaries
re.DOTALL # . matches newline too
re.VERBOSE # allow whitespace & comments in pattern
# Combine flags with |
pattern = re.compile(r"""
^ # start of line
(\w+) # capture word
\s+ # whitespace
(\d+) # capture number
""", re.VERBOSE | re.MULTILINE)
# Named groups (more readable)
m = re.match(r"(?P<year>\d{4})-(?P<month>\d{2})", "2024-01")
print(m.group("year")) # 2024
print(m.group("month")) # 01
print(m.groupdict()) # {'year': '2024', 'month': '01'}
# Lookahead/lookbehind
re.findall(r"\d+(?= dollars)", "100 dollars, 200 euros")
# ['100'] (positive lookahead)
re.findall(r"(?<=\$)\d+", "$100 and $200")
# ['100', '200'] (positive lookbehind)Async & Concurrency
asyncio Basics
asyncio is Python's async I/O framework. 'async def' defines a coroutine; 'await' suspends until a result is ready. asyncio.run() starts the event loop. asyncio.gather() runs coroutines concurrently. Coroutines enable high-concurrency I/O without threads.
import asyncio
# Define a coroutine
async def greet(name, delay):
await asyncio.sleep(delay) # non-blocking sleep
return f"Hello, {name}!"
# Run a coroutine
async def main():
result = await greet("Alice", 1)
print(result)
# Run the event loop
asyncio.run(main())
# Concurrent execution with gather
async def main():
# Run coroutines concurrently
results = await asyncio.gather(
greet("Alice", 2),
greet("Bob", 1),
greet("Carol", 3)
)
print(results) # all complete after 3 seconds (max delay)
asyncio.run(main())
# asyncio.create_task - schedule without awaiting immediately
async def main():
task = asyncio.create_task(greet("Alice", 1))
# do other work here
result = await task # await when needed
print(result)Async HTTP & Timeouts
asyncio.timeout() (Python 3.11+) cancels operations that take too long. For HTTP, use aiohttp (async) instead of requests (sync). asyncio.Queue enables producer-consumer patterns. Async is ideal for I/O-bound work (network, disk) — not CPU-bound work.
import asyncio
# Async timeout
async def fetch_with_timeout(url, timeout=5):
try:
async with asyncio.timeout(timeout):
# simulate async operation
await asyncio.sleep(2)
return f"Data from {url}"
except asyncio.TimeoutError:
return "Request timed out"
# Using aiohttp (third-party) for HTTP
# import aiohttp
#
# async def fetch(url):
# async with aiohttp.ClientSession() as session:
# async with session.get(url) as response:
# return await response.text()
# Producer-consumer pattern
async def producer(queue):
for i in range(5):
await asyncio.sleep(0.1)
await queue.put(f"item-{i}")
await queue.put(None) # sentinel
async def consumer(queue):
while True:
item = await queue.get()
if item is None:
break
print(f"Processed: {item}")
queue.task_done()
async def main():
queue = asyncio.Queue()
await asyncio.gather(producer(queue), consumer(queue))
asyncio.run(main())Threading
Threading is for I/O-bound concurrency (network, file I/O). Python's GIL prevents true parallel CPU execution in threads. Use Lock to protect shared state from race conditions. For CPU-bound work, use multiprocessing instead. Threads share memory; processes don't.
import threading
import time
# Basic threading
def worker(name, delay):
print(f"Worker {name} starting")
time.sleep(delay)
print(f"Worker {name} done")
# Create and start threads
t1 = threading.Thread(target=worker, args=("A", 2))
t2 = threading.Thread(target=worker, args=("B", 1))
t1.start()
t2.start()
# Wait for threads to complete
t1.join()
t2.join()
print("All done")
# Thread with Lock (for shared state)
counter = 0
lock = threading.Lock()
def increment():
global counter
for _ in range(100000):
with lock: # acquire/release lock
counter += 1
threads = [threading.Thread(target=increment) for _ in range(5)]
for t in threads: t.start()
for t in threads: t.join()
print(f"Counter: {counter}") # 500000 (correct with lock)Multiprocessing
Multiprocessing bypasses the GIL for true CPU parallelism — each process has its own Python interpreter. Use Pool for parallel map operations. concurrent.futures provides a unified API for both threads and processes. Always guard with if __name__ == '__main__' on Windows.
from multiprocessing import Process, Pool, Queue
import os
# Basic process
def worker(name):
print(f"Process {name} PID: {os.getpid()}")
if __name__ == "__main__":
p = Process(target=worker, args=("A",))
p.start()
p.join()
# Process pool for parallel work
def square(x):
return x * x
if __name__ == "__main__":
with Pool(4) as pool: # 4 worker processes
results = pool.map(square, range(10))
print(results) # [0, 1, 4, 9, ..., 81]
# Asynchronous map
with Pool(4) as pool:
result = pool.map_async(square, range(10))
print(result.get()) # blocks until done
# concurrent.futures (higher-level API)
from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor
with ProcessPoolExecutor() as executor:
results = list(executor.map(square, range(10)))
with ThreadPoolExecutor() as executor:
futures = [executor.submit(square, i) for i in range(10)]
results = [f.result() for f in futures]Decorators
Basic Decorator
Decorators wrap a function to extend or modify its behavior without changing the original source. The @syntax is syntactic sugar for assigning the result of the decorator call back to the function name. Use *args, **kwargs in the wrapper so it works with any signature.
# A decorator is a function that takes a function and returns a new function
def uppercase_result(func):
def wrapper(*args, **kwargs):
result = func(*args, **kwargs)
return result.upper()
return wrapper
@uppercase_result
def greet(name):
return f"hello, {name}"
print(greet("world")) # HELLO, WORLD
# @uppercase_result is sugar for: greet = uppercase_result(greet)functools.wraps (Preserve Metadata)
Without @wraps, the wrapped function loses its original __name__, __doc__, and signature — debugging tools and help() show 'wrapper' instead. Always use @functools.wraps(func) inside decorators to preserve metadata. This is a near-universal best practice.
from functools import wraps
def log_calls(func):
@wraps(func) # copies __name__, __doc__, __module__
def wrapper(*args, **kwargs):
print(f"Calling {func.__name__}({args}, {kwargs})")
return func(*args, **kwargs)
return wrapper
@log_calls
def add(a, b):
"""Add two numbers."""
return a + b
print(add.__name__) # 'add' (not 'wrapper')
print(add.__doc__) # 'Add two numbers.'
help(add) # shows original docstringDecorator with Arguments
When a decorator takes arguments, you need three levels of nesting: the factory (takes args), the decorator (takes the function), and the wrapper (takes call args). @repeat(3) calls repeat(3) first, which returns the decorator, which is then applied to the function.
# A decorator factory: returns the actual decorator
def repeat(times):
def decorator(func):
@wraps(func)
def wrapper(*args, **kwargs):
result = None
for _ in range(times):
result = func(*args, **kwargs)
return result
return wrapper
return decorator
@repeat(times=3)
def say_hi(name):
print(f"Hi, {name}!")
say_hi("Alice")
# Hi, Alice! (printed 3 times)Class-Based Decorator
Class decorators use __init__ to store the function and __call__ to intercept invocations. They're ideal when the decorator needs to maintain state (like a call counter or cache). The class instance replaces the function, so calling it triggers __call__.
class CountCalls:
def __init__(self, func):
self.func = func
self.count = 0
wraps(func)(self) # preserve metadata
def __call__(self, *args, **kwargs):
self.count += 1
print(f"{self.func.__name__} called {self.count} times")
return self.func(*args, **kwargs)
@CountCalls
def say_hello():
print("Hello!")
say_hello() # count=1
say_hello() # count=2
say_hello() # count=3
print(say_hello.count) # 3Built-in Decorators (@property, @staticmethod, @classmethod)
@property turns a method into a computed attribute (accessed without parentheses). @classmethod receives the class as the first argument — perfect for alternative constructors. @staticmethod receives no implicit first argument — just a function that happens to live in the class namespace. Together they form the backbone of Pythonic OOP.
class Circle:
pi = 3.14159
def __init__(self, radius):
self._radius = radius
@property
def area(self):
return Circle.pi * self._radius ** 2
@property
def radius(self):
return self._radius
@radius.setter
def radius(self, value):
if value < 0:
raise ValueError("Radius cannot be negative")
self._radius = value
@classmethod
def from_diameter(cls, diameter):
return cls(diameter / 2)
@staticmethod
def is_valid_radius(r):
return r >= 0
c = Circle(5)
print(c.area) # 78.54 (no parentheses!)
c.radius = 10 # uses the setter
c2 = Circle.from_diameter(20) # alternative constructorStacked Decorators
When stacking decorators, they are applied bottom-up (the one closest to the function runs first) but execute top-down at call time. So @bold wraps @italic wraps greet. The result is nested like onion layers. Order matters — reversing them changes the output nesting.
from functools import wraps
def bold(func):
@wraps(func)
def wrapper(*args, **kwargs):
return f"<b>{func(*args, **kwargs)}</b>"
return wrapper
def italic(func):
@wraps(func)
def wrapper(*args, **kwargs):
return f"<i>{func(*args, **kwargs)}</i>"
return wrapper
@bold
@italic
def greet(name):
return f"Hello, {name}"
print(greet("World"))
# <b><i>Hello, World</i></b>
# Applied bottom-up: italic first, then boldGenerators & Iterators
Generator Functions (yield)
Generators produce values lazily using yield — they pause execution after each yield and resume when next() is called. This makes them memory-efficient for large or infinite sequences since only one value exists in memory at a time. Once exhausted, a generator cannot be reused.
# A generator uses 'yield' to produce values lazily, one at a time
def count_up_to(n):
count = 1
while count <= n:
yield count
count += 1
gen = count_up_to(5)
print(next(gen)) # 1
print(next(gen)) # 2
print(list(gen)) # [3, 4, 5] (exhausts the rest)
# Generators are memory-efficient: they don't build the whole list
def fibonacci():
a, b = 0, 1
while True:
yield a
a, b = b, a + b
fib = fibonacci()
print([next(fib) for _ in range(10)])
# [0, 1, 1, 2, 3, 5, 8, 13, 21, 34]Generator Expressions
Generator expressions are the lazy equivalent of list comprehensions — use parentheses instead of brackets. They use constant memory regardless of size, making them ideal for sum(), max(), any(), or feeding into other iterators. Prefer them over list comprehensions when you don't need random access.
# Like list comprehensions, but lazy (uses parentheses)
squares_list = [x ** 2 for x in range(10)] # builds full list
squares_gen = (x ** 2 for x in range(10)) # lazy generator
print(squares_gen) # <generator object>
print(next(squares_gen)) # 0
print(next(squares_gen)) # 1
# Memory comparison
import sys
print(sys.getsizeof([x for x in range(10000)])) # ~87616 bytes
print(sys.getsizeof((x for x in range(10000)))) # ~200 bytes (constant!)
# Use in sum(), list(), any() etc.
total = sum(x ** 2 for x in range(100)) # no extra list created
print(total) # 328350Iterator Protocol (__iter__, __next__)
The iterator protocol requires __iter__ (returns an iterator) and __next__ (returns the next value or raises StopIteration). Iterables can be looped over; iterators produce values one at a time. For reusable iterables, separate the iterable (returns a new iterator) from the iterator (holds the state).
class Range2:
"""A custom iterator that yields even numbers."""
def __init__(self, start, end):
self.current = start
self.end = end
def __iter__(self):
return self # the object is its own iterator
def __next__(self):
if self.current >= self.end:
raise StopIteration
value = self.current
self.current += 2
return value
r = Range2(0, 10)
for num in r:
print(num) # 0, 2, 4, 6, 8
# An iterable returns a fresh iterator each time __iter__ is called.
# An iterator returns itself and raises StopIteration when exhausted.send(), throw(), close()
Advanced generator methods enable two-way communication: send() passes a value into the generator (becomes the result of yield), throw() injects an exception at the yield point, and close() terminates the generator. You must 'prime' the generator with next() before sending. These power coroutines and async frameworks.
def echo():
while True:
received = yield # yield without a value, receives via send()
print(f"Echo: {received}")
gen = echo()
next(gen) # prime the generator (advance to first yield)
gen.send("hello") # Echo: hello
gen.send("world") # Echo: world
# throw() injects an exception at the yield point
def safe_gen():
try:
while True:
yield
except ValueError:
print("Caught ValueError inside generator")
g = safe_gen()
next(g)
g.throw(ValueError, "boom") # Caught ValueError inside generator
# close() stops the generator (raises GeneratorExit)
gen.close()Generator Pipelines
Generator pipelines chain lazy producers so data flows through stages one item at a time — each item is fully processed before the next is read. This avoids building intermediate lists and is the foundation of streaming data processing. Unix pipes work the same way conceptually.
# Chain generators to build data-processing pipelines
def numbers():
for i in range(1, 11):
yield i
def squared(seq):
for n in seq:
yield n ** 2
def evens(seq):
for n in seq:
if n % 2 == 0:
yield n
# Each stage processes one item at a time — no intermediate lists
pipeline = evens(squared(numbers()))
print(list(pipeline)) # [4, 16, 36, 64, 100]
# Equivalent with generator expressions:
result = (n for n in (x ** 2 for x in range(1, 11)) if n % 2 == 0)
print(list(result)) # [4, 16, 36, 64, 100]yield from (Delegation)
yield from delegates all yields (and send/throw/close) to a sub-iterator, flattening nested structures and composing coroutines. It's especially powerful for recursive generators — the classic use case is flattening arbitrarily nested lists. In async code, 'await' is built on the same concept.
# yield from delegates to a sub-iterator (Python 3.3+)
def flatten(nested):
for item in nested:
if isinstance(item, (list, tuple)):
yield from flatten(item) # recursive delegation
else:
yield item
data = [1, [2, 3, [4, 5]], 6, [7, [8, [9]]]]
print(list(flatten(data)))
# [1, 2, 3, 4, 5, 6, 7, 8, 9]
# yield from also forwards send()/throw() to the sub-generator,
# making it essential for coroutine composition.Context Managers
with Statement Basics
The with statement ensures resources are released (files closed, locks released, connections returned) even when exceptions occur. It calls __enter__ at the start and __exit__ at the end. Always prefer 'with' over manual try/finally for resource management — it's safer and more readable.
# 'with' guarantees cleanup even if an exception occurs
with open("data.txt", "r") as f:
content = f.read()
# f.close() is called automatically here, even if read() raised
# Without 'with' you must manually close:
f = open("data.txt", "r")
try:
content = f.read()
finally:
f.close() # easy to forget!
# Common built-in context managers:
with open("out.txt", "w") as f, open("in.txt") as g:
f.write(g.read()) # both files close automaticallyCustom Context Manager (Class)
A class-based context manager implements __enter__ (setup, returns the context object) and __exit__(exc_type, exc_val, exc_tb) (cleanup). The __exit__ arguments receive exception info if one occurred; returning True suppresses it. This pattern is ideal for complex setup/teardown like database transactions.
class Timer:
def __init__(self, label="Timer"):
self.label = label
def __enter__(self):
import time
self.start = time.perf_counter()
return self # value bound to 'as' variable
def __exit__(self, exc_type, exc_val, exc_tb):
import time
elapsed = time.perf_counter() - self.start
print(f"{self.label}: {elapsed:.4f}s")
# Return False (or None) to propagate exceptions
# Return True to suppress the exception
return False
with Timer("Processing"):
total = sum(i ** 2 for i in range(1_000_000))
# Processing: 0.1234scontextlib.contextmanager
contextlib.contextmanager turns a generator function into a context manager — code before yield is __enter__, code after yield (in finally) is __exit__. This is more concise than a class for simple cases. Yield a value to provide it to the 'as' variable. Use try/finally to guarantee cleanup.
from contextlib import contextmanager
import time
@contextmanager
def timer(label="Timer"):
start = time.perf_counter()
try:
yield # code inside the 'with' block runs here
finally:
elapsed = time.perf_counter() - start
print(f"{label}: {elapsed:.4f}s")
with timer("My task"):
sum(i ** 2 for i in range(1_000_000))
# My task: 0.1234s
# You can also yield a value to bind with 'as'
@contextmanager
def open_db(path):
db = connect(path)
try:
yield db
finally:
db.close()
with open_db("app.db") as db:
db.query("SELECT 1")Multiple Context Managers
Python 3.10+ allows parenthesized multi-line 'with' statements for cleaner syntax. For dynamic numbers of context managers, contextlib.ExitStack manages them as a group and unwinds all in reverse order. ExitStack is essential when the number of resources isn't known until runtime.
# Python 3.10+ supports parenthesized context managers
with (
open("input.txt") as fin,
open("output.txt", "w") as fout,
):
fout.write(fin.read())
# Pre-3.10: nest them or use contextlib.ExitStack
from contextlib import ExitStack
files = ["a.txt", "b.txt", "c.txt"]
with ExitStack() as stack:
handles = [stack.enter_context(open(f)) for f in files]
for h in handles:
print(h.read())contextlib Utilities (suppress, redirect)
contextlib.suppress replaces try/except/pass for expected exceptions — much more readable. redirect_stdout/redirect_stderr capture output that would otherwise go to the console, useful for testing or logging. These utilities avoid boilerplate and make intent explicit.
from contextlib import suppress, redirect_stdout, redirect_stderr
import io, warnings
# suppress: ignore specific exceptions (cleaner than try/except/pass)
with suppress(FileNotFoundError):
os.remove("temp.txt") # no error if file doesn't exist
# redirect_stdout: capture print output
buffer = io.StringIO()
with redirect_stdout(buffer):
print("This goes to the buffer, not console")
captured = buffer.getvalue()
# redirect_stderr: capture error/warning output
err_buf = io.StringIO()
with redirect_stderr(err_buf):
warnings.warn("a warning")
print(err_buf.getvalue()) # the warning text
# Also: contextlib.chdir (3.11+) to temporarily change directory
# from contextlib import chdir
# with chdir("/tmp"): ...Async Context Managers
Async context managers use __aenter__/__aexit__ (note the 'a' prefix) and the 'async with' statement. They're essential for managing async resources like database connections or HTTP sessions (e.g., aiohttp.ClientSession). The cleanup runs even if an await or exception occurs inside the block.
import asyncio
class AsyncDB:
async def __aenter__(self):
print("connecting...")
await asyncio.sleep(0.1) # simulate async connect
return self
async def __aexit__(self, exc_type, exc_val, exc_tb):
print("closing...")
await asyncio.sleep(0.1) # simulate async close
return False
async def query(self, sql):
return f"result of: {sql}"
async def main():
async with AsyncDB() as db:
result = await db.query("SELECT 1")
print(result)
asyncio.run(main())
# connecting...
# result of: SELECT 1
# closing...Type Hints
Basic Variable & Function Annotations
Type hints document expected types but are NOT enforced at runtime — Python remains dynamically typed. Use a static checker like mypy or pyright to catch type errors before running. The built-in generics (list[str], dict[str, int]) require Python 3.9+; older versions need typing.List, typing.Dict.
# Variable annotations (Python 3.6+)
name: str = "Alice"
age: int = 30
scores: list[float] = [95.5, 88.0, 92.3]
config: dict[str, int] = {"timeout": 30}
# Function annotations
def greet(name: str, excited: bool = False) -> str:
punctuation = "!" if excited else "."
return f"Hello, {name}{punctuation}"
print(greet("World", excited=True)) # Hello, World!
# Annotations are optional and NOT enforced at runtime
def add(a: int, b: int) -> int:
return a + b
add("2", "3") # runs fine, returns "23" (no error!)typing Module (List, Dict, Tuple, Optional)
The typing module provides generic aliases for older Python. Since 3.9, you can use built-in types directly (list[str] instead of List[str]). Optional[X] is shorthand for Union[X, None] — use it to signal that a function may return None, forcing callers to handle the None case.
from typing import List, Dict, Tuple, Set, FrozenSet
# Pre-3.9 style (still works, needed for older Python)
names: List[str] = ["Alice", "Bob"]
scores: Dict[str, int] = {"Alice": 95}
point: Tuple[int, int] = (10, 20)
mixed: Tuple[str, int, float] = ("a", 1, 2.0)
variadic: Tuple[int, ...] = (1, 2, 3) # variable-length
# Python 3.9+ built-in generics (preferred)
names: list[str] = ["Alice", "Bob"]
scores: dict[str, int] = {"Alice": 95}
point: tuple[int, int] = (10, 20)
# Optional means "could be None"
from typing import Optional
def find_user(uid: int) -> Optional[str]:
if uid == 1:
return "Alice"
return None # could also just 'return'Union & Literal Types
Union[X, Y] (or X | Y in 3.10+) means a value can be either type. Literal restricts a value to specific constants — great for string enums without the enum overhead, and for overloaded function dispatch. mypy uses Literal to narrow types and catch invalid arguments at check time.
from typing import Union, Literal, overload
# Union: value can be one of several types
def process(data: Union[str, bytes]) -> str:
if isinstance(data, bytes):
return data.decode("utf-8")
return data
# Python 3.10+ union syntax with | (preferred)
def process2(data: str | bytes) -> str:
if isinstance(data, bytes):
return data.decode("utf-8")
return data
# Literal: restrict to specific constant values
def set_mode(mode: Literal["r", "w", "a"]) -> None:
print(f"Mode set to {mode}")
set_mode("r") # OK
# set_mode("x") # mypy error: not a valid literal
# Literal for boolean-like flags
Direction = Literal["up", "down", "left", "right"]TypeVar & Generics
TypeVar creates generic type variables so functions and classes can preserve type relationships (e.g., 'returns the same type as the input'). Use bound= to constrain to a subtype, or specify constraints like TypeVar('T', int, float). Generic classes use Generic[T] as a base to become parameterized containers.
from typing import TypeVar, Generic, List
T = TypeVar("T") # a generic type variable
def first(items: List[T]) -> T:
return items[0]
# Type inference: T is bound to the argument's type
x: int = first([1, 2, 3]) # T = int
y: str = first(["a", "b", "c"]) # T = str
# Bounded TypeVar: T must be a subtype of Number
from typing import TypeVar
from numbers import Number
N = TypeVar("N", bound=Number)
def sum_all(values: list[N]) -> N:
total = values[0]
for v in values[1:]:
total = total + v
return total
# Generic class
class Stack(Generic[T]):
def __init__(self) -> None:
self._items: list[T] = []
def push(self, item: T) -> None:
self._items.append(item)
def pop(self) -> T:
return self._items.pop()
s: Stack[int] = Stack()
s.push(1)
# s.push("x") # mypy errorCallable, Type Aliases & Protocols
Callable[[int, str], bool] describes a function taking int and str, returning bool. Type aliases give descriptive names to complex types. Protocol enables structural (duck) typing — any object with the right methods satisfies the protocol, no inheritance required. This is Python's answer to interfaces.
from typing import Callable, Protocol, TypeAlias
# Callable signature: Callable[[ArgTypes], ReturnType]
def apply(func: Callable[[int, int], int], a: int, b: int) -> int:
return func(a, b)
apply(lambda x, y: x + y, 3, 4) # 7
# Type aliases (3.12+ uses 'type' statement; older uses assignment)
type Vector = list[float] # 3.12+
Vector2: TypeAlias = list[float] # 3.10+
def magnitude(v: Vector) -> float:
return sum(x ** 2 for x in v) ** 0.5
# Protocol: structural typing (duck typing with static checks)
class Drawable(Protocol):
def draw(self) -> None: ...
def render(obj: Drawable) -> None:
obj.draw() # any object with a draw() method works
class Circle:
def draw(self) -> None:
print("drawing circle")
render(Circle()) # OK — Circle has draw()Type Checking with mypy
mypy is the most popular static type checker for Python — it analyzes type hints without running the code. It catches None-handling bugs, wrong argument types, and missing returns. Start with gradual typing: add hints to new code and run mypy in CI. Use --strict for new projects to enforce comprehensive annotations.
# Save as example.py, then run: mypy example.py
from typing import Optional
def divide(a: float, b: float) -> Optional[float]:
if b == 0:
return None
return a / b
result = divide(10, 0)
# Without checking, this would crash at runtime:
# print(result + 1) # TypeError: NoneType + int
# mypy catches it:
# error: Unsupported operand types for + ("None" and "int")
# fix: check for None first
if result is not None:
print(result + 1)
# Run strict mode for maximum safety:
# mypy --strict example.py
# Common strict flags: --disallow-untyped-defs, --no-implicit-optionalData Classes
Basic @dataclass
@dataclass auto-generates __init__, __repr__, and __eq__ based on annotated fields — eliminating boilerplate for data-holding classes. It's ideal for value objects, configs, DTOs, and records. Available since Python 3.7. Fields must have type annotations; the annotation defines the field.
from dataclasses import dataclass
@dataclass
class Point:
x: float
y: float
p1 = Point(1.0, 2.0)
p2 = Point(1.0, 2.0)
print(p1) # Point(x=1.0, y=2.0) — auto __repr__
print(p1 == p2) # True — auto __eq__ (compares fields)
print(p1.x) # 1.0
# Without @dataclass you'd write all this boilerplate:
# class Point:
# def __init__(self, x, y): self.x = x; self.y = y
# def __repr__(self): ...
# def __eq__(self, other): ...Default Values & default_factory
Mutable defaults (lists, dicts, sets) must use field(default_factory=list) — using [] directly would share one list across all instances, a classic bug. default_factory is called once per instance to create a fresh object. Simple immutable defaults (int, str, bool, None) can be assigned directly.
from dataclasses import dataclass, field
@dataclass
class Student:
name: str
grade: str = "A" # simple default
tags: list[str] = field(default_factory=list) # mutable default!
scores: dict[str, int] = field(default_factory=dict)
s = Student("Alice")
print(s) # Student(name='Alice', grade='A', tags=[], scores={})
s.tags.append("honors")
s2 = Student("Bob")
print(s2.tags) # [] — each instance gets its own list
# NEVER use [] or {} as a direct default — all instances would share
# the same mutable object (classic Python pitfall).frozen, eq & order
frozen=True makes a dataclass immutable — fields can't be reassigned, and the instance becomes hashable (usable as dict keys or set members). order=True adds comparison methods for sorting. Combine frozen=True with order=True for immutable, sortable value types like coordinates, colors, or versions.
from dataclasses import dataclass
# frozen=True makes instances immutable (hashable, usable as dict keys)
@dataclass(frozen=True)
class Color:
r: int
g: int
b: int
c = Color(255, 0, 0)
# c.r = 128 # FrozenInstanceError!
print(hash(c)) # works — frozen dataclasses are hashable
# order=True generates __lt__, __le__, __gt__, __ge__ for sorting
@dataclass(order=True)
class Priority:
level: int
tasks = [Priority(3), Priority(1), Priority(2)]
tasks.sort()
print(tasks) # [Priority(level=1), Priority(level=2), Priority(level=3)]
# Common combo: frozen + order for immutable comparable values
@dataclass(frozen=True, order=True)
class Version:
major: int
minor: int__post_init__ & field customization
__post_init__ runs automatically after the generated __init__ — use it to compute derived fields, validate values, or perform setup. field(init=False) creates a field not in the constructor (good for computed/cached values). field(repr=False, compare=False) hides fields from repr and equality checks.
from dataclasses import dataclass, field
@dataclass
class User:
email: str
_email_normalized: str = field(init=False, repr=False)
id: int = field(default=0)
def __post_init__(self):
# runs after __init__; compute derived fields here
self._email_normalized = self.email.strip().lower()
u = User(" [email protected] ")
print(u.email) # ' [email protected] '
print(u._email_normalized) # '[email protected]'
# field(init=False) excludes a field from __init__
# field(repr=False) hides it from the repr
# field(compare=False) excludes from __eq__/__hash__
# field(metadata={...}) attaches custom metadataInheritance & slots
Dataclasses support inheritance — child fields are appended after parent fields, and you can override parent defaults. Note: a field with a default in the parent cannot be followed by a non-default field in the child. slots=True (3.10+) prevents adding arbitrary attributes and significantly reduces memory per instance — ideal for millions of small objects.
from dataclasses import dataclass
@dataclass
class Animal:
name: str
sound: str = "..."
@dataclass
class Dog(Animal):
breed: str = "unknown"
sound: str = "Woof" # override parent default
d = Dog("Rex", breed="Labrador")
print(d) # Dog(name='Rex', sound='Woof', breed='Labrador')
# Python 3.10+: slots=True saves memory (no __dict__)
@dataclass(slots=True)
class Pixel:
r: int
g: int
b: int
p = Pixel(0, 128, 255)
# p.new_field = 1 # AttributeError — slots prevent arbitrary attrs
# Saves ~40-50% memory vs regular dataclass for many instancesCollections & Itertools
namedtuple
namedtuple creates tuple subclasses with named fields — as memory-efficient as tuples but far more readable. They're immutable, so use _replace() to create modified copies. Prefer typing.NamedTuple for new code since it supports type annotations and default values. Great for returning multiple values from functions.
from collections import namedtuple
# Lightweight immutable class with named fields
Point = namedtuple("Point", ["x", "y"])
p = Point(3, 4)
print(p.x, p.y) # 3 4 — access by name
print(p[0], p[1]) # 3 4 — also by index
print(p._asdict()) # {'x': 3, 'y': 4}
# More memory-efficient than a full class
# Use _replace to create a modified copy (immutable!)
p2 = p._replace(x=10)
print(p2) # Point(x=10, y=4)
# Python 3.6+ typing.NamedTuple for type hints:
from typing import NamedTuple
class Point3D(NamedTuple):
x: float
y: float
z: float = 0.0Counter
Counter is a dict subclass for counting hashable objects — perfect for frequency analysis, histograms, and voting. most_common(n) returns the top n items. Missing keys return 0 instead of raising KeyError. Counter supports +, -, &, | for set-like arithmetic on counts.
from collections import Counter
# Count hashable items
words = "the cat sat on the mat the cat".split()
c = Counter(words)
print(c) # Counter({'the': 3, 'cat': 2, 'sat': 1, 'on': 1, 'mat': 1})
# Most common elements
print(c.most_common(2)) # [('the', 3), ('cat', 2)]
# Arithmetic on counters
c1 = Counter(a=3, b=1)
c2 = Counter(a=1, b=2)
print(c1 + c2) # Counter({'a': 4, 'b': 3})
print(c1 - c2) # Counter({'a': 2}) (drops zero/negatives)
# Missing keys return 0 (not KeyError)
print(c["dog"]) # 0
# Update and elements
c.update(["cat", "cat"])
print(sorted(c.elements())) # ['cat','cat','cat','cat','mat','on','sat','the','the','the']defaultdict
defaultdict automatically creates missing keys with a default value from a factory function — list for grouping, int for counting, set for deduplication. This eliminates the 'if key not in dict' boilerplate. The factory is called only when a key is missing, not on every access.
from collections import defaultdict
# Group items by key without checking if key exists
words = ["apple", "banana", "avocado", "blueberry", "cherry"]
by_first = defaultdict(list)
for w in words:
by_first[w[0]].append(w)
print(dict(by_first))
# {'a': ['apple', 'avocado'], 'b': ['banana', 'blueberry'], 'c': ['cherry']}
# Counting with int (default 0)
counts = defaultdict(int)
for w in words:
counts[w[0]] += 1
print(dict(counts)) # {'a': 2, 'b': 2, 'c': 1}
# Nested defaultdicts
tree = defaultdict(lambda: defaultdict(list))
tree["2024"]["Jan"].append("event1")
# vs regular dict: avoids the key-check boilerplate
# d = {}
# for w in words:
# if w[0] not in d:
# d[w[0]] = []
# d[w[0]].append(w)OrderedDict & deque
deque provides O(1) append/pop at both ends — use it for queues, BFS, and sliding windows instead of lists (list.pop(0) is O(n)). With maxlen, deque auto-discards old items, perfect for bounded buffers. OrderedDict is less needed since 3.7 (dicts are ordered), but its move_to_end and popitem are still uniquely useful for LRU caches.
from collections import OrderedDict, deque
# deque: double-ended queue, O(1) append/pop at both ends
dq = deque([1, 2, 3], maxlen=5)
dq.appendleft(0) # deque([0, 1, 2, 3])
dq.append(4) # deque([0, 1, 2, 3, 4])
dq.append(5) # deque([1, 2, 3, 4, 5]) — oldest dropped (maxlen!)
print(dq.popleft()) # 1
print(dq) # deque([2, 3, 4, 5])
# deque is ideal for queues, BFS, sliding windows
from collections import deque
queue = deque(["task1", "task2"])
queue.append("task3")
next_task = queue.popleft() # FIFO — O(1) vs list.pop(0) which is O(n)
# OrderedDict: remembers insertion order (regular dicts do too in 3.7+,
# but OrderedDict has move_to_end and equality is order-sensitive)
od = OrderedDict([("a", 1), ("b", 2)])
od.move_to_end("a") # move to last
print(list(od)) # ['b', 'a']
od.popitem(last=False) # pop first item (FIFO)itertools: chain, product, combinations, permutations
itertools provides fast, memory-efficient tools for combinatorics. chain flattens iterables lazily. product gives Cartesian products (replaces nested for loops). combinations/permutations generate selections without building the full list — essential for large or infinite inputs. All return iterators, so wrap in list() to view.
from itertools import chain, product, combinations, permutations
# chain: flatten multiple iterables
list(chain([1, 2], [3, 4], [5])) # [1, 2, 3, 4, 5]
list(chain.from_iterable([[1, 2], [3, 4]])) # [1, 2, 3, 4]
# product: Cartesian product (nested loops)
list(product([1, 2], ["a", "b"]))
# [(1,'a'), (1,'b'), (2,'a'), (2,'b')]
list(product("AB", repeat=2)) # [('A','A'),('A','B'),('B','A'),('B','B')]
# combinations: unordered selections (no repeats)
list(combinations("ABC", 2)) # [('A','B'),('A','C'),('B','C')]
list(combinations("AAA", 2)) # [('A','A'),('A','A'),('A','A')]
# permutations: ordered arrangements
list(permutations("ABC", 2)) # [('A','B'),('A','C'),('B','A'),('B','C'),('C','A'),('C','B')]
# combinations_with_replacement: allow picking same element
from itertools import combinations_with_replacement
list(combinations_with_replacement("AB", 2)) # [('A','A'),('A','B'),('B','B')]itertools: groupby, accumulate, starmap
groupby groups consecutive elements sharing a key — sort by the key first or you'll get multiple groups for the same key. accumulate produces running totals/products. islice, takewhile, and dropwhile are lazy alternatives to slicing and filtering that work on any iterator, including infinite ones.
from itertools import groupby, accumulate, starmap, islice, takewhile, dropwhile
# groupby: group consecutive items by a key (sort first!)
data = [("A", 1), ("A", 2), ("B", 3), ("B", 4), ("A", 5)]
data.sort(key=lambda x: x[0]) # MUST sort by key first
for key, group in groupby(data, key=lambda x: x[0]):
print(key, list(group))
# A [('A',1),('A',2),('A',5)]
# B [('B',3),('B',4)]
# accumulate: running aggregate (sum by default)
list(accumulate([1, 2, 3, 4])) # [1, 3, 6, 10]
import operator
list(accumulate([1, 2, 3, 4], operator.mul)) # [1, 2, 6, 24]
# starmap: unpack args from tuples before calling
list(starmap(pow, [(2, 3), (3, 2), (10, 3)])) # [8, 9, 1000]
# islice: slice an iterator (doesn't support negative indices)
list(islice(range(100), 5, 10)) # [5, 6, 7, 8, 9]
# takewhile / dropwhile: filter by predicate
list(takewhile(lambda x: x < 5, [1, 4, 6, 3, 8])) # [1, 4]
list(dropwhile(lambda x: x < 5, [1, 4, 6, 3, 8])) # [6, 3, 8]functools: lru_cache, partial, reduce
lru_cache memoizes results — dramatic speedups for recursive or expensive pure functions; cache_info() shows hit/miss stats. partial pre-fills arguments to create specialized callables. reduce applies a function cumulatively (though sum(), any(), all() often replace it). cached_property computes once then caches on the instance.
from functools import lru_cache, partial, reduce
import operator
# lru_cache: memoize function results (Least Recently Used)
@lru_cache(maxsize=128)
def fib(n):
if n < 2:
return n
return fib(n - 1) + fib(n - 2)
print(fib(100)) # instant (without cache: impossibly slow)
print(fib.cache_info()) # CacheInfo(hits=98, misses=101, ...)
# partial: fix some arguments, create a new callable
def power(base, exponent):
return base ** exponent
square = partial(power, exponent=2)
cube = partial(power, exponent=3)
print(square(5)) # 25
print(cube(3)) # 27
# reduce: cumulatively apply a function, reducing to one value
product = reduce(operator.mul, [1, 2, 3, 4]) # 24
# Equivalent: ((1*2)*3)*4
# Python 3.8+: cached_property for lazy computed attributes
from functools import cached_property
class Data:
@cached_property
def expensive(self):
print("computing...")
return [i ** 2 for i in range(1000000)]JSON & CSV Processing
json.dumps & json.loads
json.dumps() (dump string) serializes a Python object to a JSON string; json.loads() (load string) parses JSON back. Use indent for readability, ensure_ascii=False to keep Unicode characters readable, and sort_keys for deterministic output. JSON keys must be strings — int keys become strings.
import json
# Serialize Python object to JSON string
data = {"name": "Alice", "age": 30, "scores": [95, 88, 92]}
json_str = json.dumps(data)
print(json_str) # {"name": "Alice", "age": 30, "scores": [95, 88, 92]}
# Pretty-print with indent
print(json.dumps(data, indent=2))
# {
# "name": "Alice",
# "age": 30,
# "scores": [95, 88, 92]
# }
# Parse JSON string to Python object
parsed = json.loads(json_str)
print(parsed["name"]) # Alice
print(type(parsed["scores"])) # <class 'list'>
# Sort keys, handle non-ASCII
print(json.dumps({"name": "Zoë"}, ensure_ascii=False, sort_keys=True))Reading & Writing JSON Files
json.dump() writes directly to a file object; json.load() reads from one. Always specify encoding='utf-8' for portability. Remember the type mapping: JSON objects become dicts, arrays become lists, and numbers become int or float. Datetimes, sets, and custom objects are NOT JSON-serializable by default.
import json
data = {"users": [{"id": 1, "name": "Alice"}, {"id": 2, "name": "Bob"}]}
# Write to file
with open("data.json", "w", encoding="utf-8") as f:
json.dump(data, f, indent=2, ensure_ascii=False)
# Read from file
with open("data.json", "r", encoding="utf-8") as f:
loaded = json.load(f)
print(loaded["users"][0]["name"]) # Alice
# Type conversions to remember:
# JSON object <-> Python dict
# JSON array <-> Python list
# JSON string <-> Python str
# JSON number <-> Python int/float
# JSON true/false <-> Python True/False
# JSON null <-> Python NoneCustom JSON Encoding (datetime, custom objects)
The json module can't serialize datetime, set, or custom classes by default. Provide a default function (called for un-serializable objects) or a JSONEncoder subclass. For round-tripping, pair a custom encoder with an object_hook in loads() to reconstruct the original types. This is how ORMs and ORMs serialize model objects.
import json
from datetime import datetime
# Default behavior: TypeError on non-serializable types
# json.dumps({"now": datetime.now()}) # TypeError!
# Solution 1: default function for unknown types
def default_encoder(obj):
if isinstance(obj, datetime):
return obj.isoformat()
if isinstance(obj, set):
return sorted(obj)
raise TypeError(f"Cannot serialize {type(obj)}")
data = {"now": datetime.now(), "tags": {"a", "b"}}
print(json.dumps(data, default=default_encoder))
# Solution 2: custom JSONEncoder subclass
class MyEncoder(json.JSONEncoder):
def default(self, obj):
if isinstance(obj, datetime):
return {"__datetime__": obj.isoformat()}
return super().default(obj)
print(json.dumps(data, cls=MyEncoder))
# Decoding with object_hook
def decoder(dct):
if "__datetime__" in dct:
return datetime.fromisoformat(dct["__datetime__"])
return dct
json.loads(json.dumps(data, cls=MyEncoder), object_hook=decoder)Reading CSV Files
Always open CSV files with newline='' to avoid blank-row issues on Windows. csv.reader returns lists; csv.DictReader returns dicts keyed by the header row. The csv module handles quoting, embedded commas, and newlines correctly — never split CSV lines manually with line.split(','). Use Sniffer to auto-detect delimiters.
import csv
# Basic reader: each row is a list of strings
with open("data.csv", newline="", encoding="utf-8") as f:
reader = csv.reader(f)
for row in reader:
print(row) # ['name', 'age', 'city']
# DictReader: each row is a dict keyed by header
with open("data.csv", newline="") as f:
reader = csv.DictReader(f)
for row in reader:
print(row["name"], row["age"]) # access by column name
# Handle different delimiters and quoting
with open("data.tsv", newline="") as f:
reader = csv.reader(f, delimiter="\t", quotechar='"')
for row in reader:
print(row)
# Sniffer to auto-detect format
with open("unknown.csv", newline="") as f:
sample = f.read(1024)
dialect = csv.Sniffer().sniff(sample)
f.seek(0)
reader = csv.reader(f, dialect)Writing CSV Files
csv.writer writes lists; csv.DictWriter writes dicts with a fixed set of fieldnames. Always use newline='' when opening the file. The quoting parameter controls when fields are quoted — QUOTE_MINIMAL (default) only quotes when needed, QUOTE_ALL quotes everything, useful for strict parsers.
import csv
rows = [
["name", "age", "city"],
["Alice", 30, "NYC"],
["Bob", 25, "LA"],
]
# Basic writer
with open("out.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.writer(f)
writer.writerows(rows) # write multiple rows
# DictWriter: write from dicts
with open("out.csv", "w", newline="") as f:
fieldnames = ["name", "age", "city"]
writer = csv.DictWriter(f, fieldnames=fieldnames)
writer.writeheader()
writer.writerow({"name": "Alice", "age": 30, "city": "NYC"})
writer.writerow({"name": "Bob", "age": 25, "city": "LA"})
# Control quoting: QUOTE_MINIMAL (default), QUOTE_ALL, QUOTE_NONNUMERIC
writer = csv.writer(f, quoting=csv.QUOTE_ALL)
# QUOTE_ALL wraps every field in quotes: "Alice","30","NYC"JSON Lines (NDJSON) & Streaming
JSON Lines (NDJSON) puts one JSON object per line — ideal for logs, event streams, and append-only files because you can process each line independently. For huge single JSON documents, use the ijson library to stream-parse without loading the entire file into memory. NDJSON is the standard for many data pipelines.
import json
# JSON Lines: one JSON object per line (great for logs, big data)
records = [{"id": 1, "msg": "first"}, {"id": 2, "msg": "second"}]
# Write NDJSON
with open("logs.jsonl", "w") as f:
for rec in records:
f.write(json.dumps(rec) + "\n")
# Read NDJSON line by line (memory-efficient for huge files)
with open("logs.jsonl", "r") as f:
for line in f:
rec = json.loads(line)
print(rec["id"], rec["msg"])
# Stream large JSON arrays without loading everything into memory
# Use ijson library for streaming parsing of huge JSON files:
# import ijson
# with open("huge.json", "rb") as f:
# for item in ijson.items(f, "items.item"):
# process(item) # one item at a timeLogging & Testing
logging Basics
The logging module is the standard way to emit diagnostic output — far better than print() because you control levels, formats, and destinations. Use logging.getLogger(__name__) per module so you can tune verbosity per module. logging.exception() automatically includes the traceback. Configure basicConfig once at startup.
import logging
# Basic configuration (call once at program start)
logging.basicConfig(
level=logging.DEBUG,
format="%(asctime)s [%(levelname)s] %(name)s: %(message)s",
datefmt="%Y-%m-%d %H:%M:%S",
)
# Log levels (severity ascending)
logging.debug("Detailed debug info") # DEBUG (10)
logging.info("General information") # INFO (20)
logging.warning("Something unexpected") # WARNING (30)
logging.error("A real error occurred") # ERROR (40)
logging.critical("System is down") # CRITICAL (50)
# Logging exceptions with traceback
try:
1 / 0
except ZeroDivisionError:
logging.exception("Division failed") # includes full traceback
# Get a named logger (best practice per module)
logger = logging.getLogger(__name__)
logger.info("Module-specific log")Logging to File & Multiple Handlers
Handlers route log records to destinations — console, files, network, email. RotatingFileHandler caps file size and keeps backups, preventing unbounded log growth. Each handler can have its own level and format (e.g., detailed logs to file, concise logs to console). TimedRotatingFileHandler rotates by time instead of size.
import logging
from logging.handlers import RotatingFileHandler
logger = logging.getLogger("myapp")
logger.setLevel(logging.DEBUG)
# Console handler (INFO and above)
console = logging.StreamHandler()
console.setLevel(logging.INFO)
console.setFormatter(logging.Formatter("%(levelname)s: %(message)s"))
# Rotating file handler (DEBUG and above, max 5MB x 3 backups)
file_handler = RotatingFileHandler(
"app.log", maxBytes=5_000_000, backupCount=3, encoding="utf-8"
)
file_handler.setLevel(logging.DEBUG)
file_handler.setFormatter(
logging.Formatter("%(asctime)s [%(levelname)s] %(name)s: %(message)s")
)
logger.addHandler(console)
logger.addHandler(file_handler)
logger.debug("debug to file only")
logger.info("info to both console and file")
logger.error("error everywhere")unittest Basics
unittest is Python's built-in test framework (xUnit style). Tests live in classes inheriting TestCase. setUp/tearDown run before/after each test for isolation. Common assertions: assertEqual, assertTrue, assertRaises, assertIn. Run with python -m unittest for auto-discovery of test_*.py files.
import unittest
def add(a, b):
return a + b
def divide(a, b):
if b == 0:
raise ValueError("Cannot divide by zero")
return a / b
class TestMath(unittest.TestCase):
def setUp(self):
# runs before each test method
self.data = [1, 2, 3]
def tearDown(self):
# runs after each test method
pass
def test_add(self):
self.assertEqual(add(1, 2), 3)
self.assertEqual(add(-1, 1), 0)
def test_add_types(self):
self.assertEqual(add("a", "b"), "ab")
def test_divide_by_zero(self):
with self.assertRaises(ValueError):
divide(1, 0)
def test_membership(self):
self.assertIn(2, self.data)
self.assertTrue(3 in self.data)
if __name__ == "__main__":
unittest.main()
# Run: python -m unittest test_file.py -vpytest Basics
pytest is the most popular Python testing tool — plain assert statements give rich failure reports, no boilerplate classes needed. pytest.raises checks exceptions with optional regex matching. pytest.approx handles float comparison imprecision. Install with pip install pytest and run with pytest -v for verbose output.
# test_math.py — pytest is simpler and more powerful than unittest
# Install: pip install pytest
# Run: pytest -v
def add(a, b):
return a + b
def divide(a, b):
if b == 0:
raise ValueError("Cannot divide by zero")
return a / b
# Plain functions, no classes required
def test_add():
assert add(1, 2) == 3
assert add(-1, 1) == 0
def test_add_strings():
assert add("hello", " world") == "hello world"
# Testing exceptions with pytest.raises
import pytest
def test_divide_by_zero():
with pytest.raises(ValueError, match="Cannot divide by zero"):
divide(1, 0)
# Approximate float comparison
def test_float():
assert 0.1 + 0.2 == pytest.approx(0.3)pytest Fixtures
Fixtures are pytest's dependency injection — they provide setup data, mock objects, or resources to tests via parameter names. yield-based fixtures handle both setup (before yield) and teardown (after yield). Scopes control reuse: 'session' creates the fixture once for the whole run, 'module' once per file, 'function' (default) once per test.
import pytest
# A fixture provides setup data/resources to tests
@pytest.fixture
def sample_list():
return [1, 2, 3, 4, 5]
# Use fixtures by passing their name as a parameter
def test_length(sample_list):
assert len(sample_list) == 5
def test_sum(sample_list):
assert sum(sample_list) == 15
# Fixture with setup AND teardown (yield)
@pytest.fixture
def db_connection():
print("\n[setup] connecting to DB")
conn = {"connected": True}
yield conn # test runs here; value passed to test
print("\n[teardown] closing DB")
conn["connected"] = False
def test_db(db_connection):
assert db_connection["connected"] is True
# Fixture scopes: function (default), class, module, session
@pytest.fixture(scope="session")
def expensive_resource():
return load_large_dataset() # created once per test sessionpytest parametrize & mocking
parametrize runs a single test function across multiple input sets — eliminates copy-paste test code and gives clear per-case output. unittest.mock.patch replaces functions/objects with mocks for isolated testing. assert_called_once_with verifies the mock was used correctly. @pytest.mark.skip and xfail handle incomplete tests gracefully.
import pytest
from unittest.mock import patch, MagicMock
# parametrize: run one test with multiple inputs
@pytest.mark.parametrize("a, b, expected", [
(1, 2, 3),
(-1, 1, 0),
(0, 0, 0),
(100, 200, 300),
])
def test_add_many(a, b, expected):
assert add(a, b) == expected
# parametrize with IDs for readable output
@pytest.mark.parametrize("x", [1, 2, 3], ids=["one", "two", "three"])
def test_ids(x):
assert x > 0
# Mocking: replace external dependencies
def fetch_user(uid):
# imagine this calls a real API
return {"id": uid, "name": "real_user"}
@patch("__main__.fetch_user")
def test_with_mock(mock_fetch):
mock_fetch.return_value = {"id": 1, "name": "mocked"}
result = fetch_user(1)
assert result["name"] == "mocked"
mock_fetch.assert_called_once_with(1)
# Skip and expected failure
@pytest.mark.skip(reason="not implemented yet")
def test_future():
pass
@pytest.mark.xfail(reason="known bug #42")
def test_known_bug():
assert 1 == 2Network Programming
TCP Server
Creates a TCP server using the socket module. bind associates the socket with an address, listen sets the backlog queue, accept blocks until a client connects. Always close connections to free file descriptors.
import socket
server = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
server.bind(('localhost', 8080))
server.listen(5)
conn, addr = server.accept()
data = conn.recv(1024)
conn.sendall(b'Hello')
conn.close()TCP Client
Creates a TCP client that connects to a server. connect establishes the connection, sendall sends all bytes, recv reads up to the specified bytes. Use encode/decode for string-to-bytes conversion.
import socket
client = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
client.connect(('localhost', 8080))
client.sendall(b'Hello Server')
response = client.recv(1024)
print(response.decode())
client.close()UDP Socket
UDP is connectionless: no handshake, no guaranteed delivery. recvfrom returns both data and sender address. Use SOCK_DGRAM for UDP. Ideal for DNS, gaming, and real-time streaming.
import socket
sock = socket.socket(socket.AF_INET, socket.SOCK_DGRAM)
sock.bind(('localhost', 9090))
data, addr = sock.recvfrom(1024)
print(f"From {addr}: {data.decode()}")
sock.sendto(b'Reply', addr)HTTP Server
The http.server module provides a simple HTTP server. Subclass BaseHTTPRequestHandler and override do_GET, do_POST. Use for development only; use gunicorn for production.
from http.server import HTTPServer, BaseHTTPRequestHandler
class Handler(BaseHTTPRequestHandler):
def do_GET(self):
self.send_response(200)
self.send_header('Content-Type', 'text/html')
self.end_headers()
self.wfile.write(b'<h1>Hello</h1>')
HTTPServer(('localhost', 8000), Handler).serve_forever()Socket Timeout
settimeout sets a timeout for all socket operations. If an operation exceeds the timeout, a socket.timeout exception is raised. Use try/finally to ensure cleanup.
import socket
sock = socket.socket(socket.AF_INET, socket.SOCK_STREAM)
sock.settimeout(5.0)
try:
sock.connect(('example.com', 80))
data = sock.recv(1024)
except socket.timeout:
print('Connection timed out')
finally:
sock.close()Database (SQLite)
Create Table
sqlite3 is built into Python. connect creates or opens a database file. CREATE TABLE IF NOT EXISTS prevents errors if the table exists. Always call commit to save changes.
import sqlite3
conn = sqlite3.connect('example.db')
cursor = conn.cursor()
cursor.execute('''CREATE TABLE IF NOT EXISTS users (
id INTEGER PRIMARY KEY AUTOINCREMENT,
name TEXT NOT NULL, email TEXT UNIQUE, age INTEGER)''')
conn.commit()Insert Data
Always use parameterized queries (? placeholders) to prevent SQL injection. lastrowid returns the auto-incremented ID. Never use string formatting for SQL values.
cursor.execute(
'INSERT INTO users (name, email, age) VALUES (?, ?, ?)',
('Alice', '[email protected]', 30))
conn.commit()
print(f"ID: {cursor.lastrowid}")Query Data
fetchall returns all matching rows as a list of tuples. fetchone returns a single row or None. For large result sets, iterate over the cursor directly.
cursor.execute('SELECT * FROM users WHERE age > ?', (25,))
rows = cursor.fetchall()
for row in rows:
print(row)
cursor.execute('SELECT * FROM users WHERE id = ?', (1,))
user = cursor.fetchone()Update & Delete
UPDATE modifies existing rows, DELETE removes them. rowcount indicates affected rows. Always use WHERE with DELETE. commit persists changes.
cursor.execute('UPDATE users SET age = ? WHERE name = ?', (31, 'Alice'))
cursor.execute('DELETE FROM users WHERE age < ?', (18,))
conn.commit()
print(f"Affected: {cursor.rowcount} rows")Row Factory
Using conn as a context manager auto-commits on success and rolls back on exception. row_factory = sqlite3.Row allows accessing columns by name.
conn = sqlite3.connect('example.db')
conn.row_factory = sqlite3.Row
with conn:
conn.execute('INSERT INTO users (name, email) VALUES (?, ?)',
('Bob', '[email protected]'))
for row in conn.execute('SELECT * FROM users'):
print(row['name'], row['email'])Web Scraping
BeautifulSoup Basics
requests fetches HTML content, BeautifulSoup parses it. html.parser is built-in; lxml is faster. Always check response.status_code before parsing.
import requests
from bs4 import BeautifulSoup
resp = requests.get('https://example.com')
soup = BeautifulSoup(resp.text, 'html.parser')
print(soup.title.string)
print(soup.find('h1').text)Find Elements
find_all returns all matching elements, find returns the first. Use class_ (with underscore). select uses CSS selectors for complex queries.
links = soup.find_all('a')
for link in links:
print(link.get('href'), link.text)
article = soup.find('div', class_='article')
items = soup.select('ul.list > li.item')Extract Tables
Tables are structured as tr (rows) containing td (data) or th (header) cells. strip removes whitespace. find_all accepts a list of tag names.
table = soup.find('table')
for row in table.find_all('tr'):
cols = row.find_all(['td', 'th'])
data = [col.text.strip() for col in cols]
print(data)Handle Pagination
Pagination is handled by following next-page links. select_one returns the first match or None. Add time.sleep between requests.
all_items = []
url = 'https://example.com/page/1'
while url:
resp = requests.get(url)
soup = BeautifulSoup(resp.text, 'html.parser')
all_items.extend([i.text for i in soup.select('.item')])
next_link = soup.select_one('a.next')
url = next_link.get('href') if next_link else NoneSave to CSV
csv.DictWriter writes dictionaries to CSV. newline prevents extra blank lines on Windows. encoding=utf-8 handles special characters.
import csv
with open('data.csv', 'w', newline='', encoding='utf-8') as f:
writer = csv.DictWriter(f, fieldnames=['name', 'price'])
writer.writeheader()
for item in scraped_data:
writer.writerow(item)Async Web (aiohttp)
HTTP Client
aiohttp provides async HTTP. ClientSession manages connection pooling. async with ensures cleanup. asyncio.run executes the coroutine.
import aiohttp, asyncio
async def fetch(url):
async with aiohttp.ClientSession() as session:
async with session.get(url) as resp:
return await resp.text()
data = asyncio.run(fetch('https://api.example.com'))Concurrent Requests
asyncio.gather runs coroutines concurrently, reducing total time. All requests share the same session. Use a semaphore to limit concurrency.
async def fetch_all(urls):
async with aiohttp.ClientSession() as session:
tasks = [session.get(url) for url in urls]
responses = await asyncio.gather(*tasks)
return [await r.text() for r in responses]Web Server
aiohttp.web creates async web servers. Routes are defined with HTTP method and path pattern. match_info extracts path parameters.
from aiohttp import web
async def handle(request):
name = request.match_info.get('name', 'World')
return web.json_response({'message': f'Hello, {name}!'})
app = web.Application()
app.add_routes([web.get('/', handle), web.get('/{name}', handle)])
web.run_app(app, port=8080)WebSocket Server
WebSockets enable bidirectional real-time communication. WebSocketResponse handles the upgrade handshake. async for iterates over messages.
async def ws_handler(request):
ws = web.WebSocketResponse()
await ws.prepare(request)
async for msg in ws:
if msg.type == aiohttp.WSMsgType.TEXT:
await ws.send_str(f'Echo: {msg.data}')
return wsSession with Cookies
ClientSession persists cookies across requests automatically. Essential for authenticated scraping. Use a single session for all requests.
async def login_and_fetch():
async with aiohttp.ClientSession() as session:
await session.post('https://example.com/login',
data={'user': 'admin', 'pass': '123'})
resp = await session.get('https://example.com/dashboard')
return await resp.text()Multiprocessing Deep Dive
Process Pool
Pool manages worker processes. map distributes work in parallel. apply_async runs a single function asynchronously. Always use if __name__ == main guard on Windows.
from multiprocessing import Pool
def square(x): return x * x
if __name__ == '__main__':
with Pool(4) as pool:
results = pool.map(square, range(10))
result = pool.apply_async(square, (100,))
print(result.get(timeout=5))Shared Memory
Value and Array create shared memory between processes. Use get_lock to synchronize access and prevent race conditions.
from multiprocessing import Value, Array
counter = Value('i', 0)
arr = Array('d', [0.0, 1.0, 2.0])
with counter.get_lock():
counter.value += 1Queue Communication
Queue enables safe communication between processes. put adds items, get retrieves them. Queue is process-safe, handling locking internally.
from multiprocessing import Process, Queue
def worker(q):
q.put('Data from worker')
if __name__ == '__main__':
q = Queue()
p = Process(target=worker, args=(q,))
p.start()
print(q.get())
p.join()Pipe
Pipe creates a two-way communication channel. send and recv transmit Python objects via pickling. Pipe is faster than Queue for point-to-point communication.
from multiprocessing import Process, Pipe
def worker(conn):
conn.send(['hello', 'world'])
msg = conn.recv()
conn.close()
if __name__ == '__main__':
parent, child = Pipe()
p = Process(target=worker, args=(child,))
p.start()
print(parent.recv())
parent.send('acknowledged')
p.join()Synchronization
Lock ensures only one process accesses a shared resource at a time. with lock acquires and releases automatically. Other primitives: RLock, Semaphore, Event.
from multiprocessing import Process, Lock
def safe_print(lock, msg):
with lock:
print(msg)
if __name__ == '__main__':
lock = Lock()
procs = [Process(target=safe_print, args=(lock, f'Task {i}'))
for i in range(5)]
for p in procs: p.start()
for p in procs: p.join()Virtual Environments Deep Dive
venv Module
venv creates isolated Python environments with their own package directories. Activation modifies PATH. Always activate before installing dependencies.
# Create
python -m venv myenv
# Activate (Linux/Mac)
source myenv/bin/activate
# Activate (Windows)
myenv\Scripts\activate
# Deactivate
deactivaterequirements.txt
requirements.txt lists project dependencies. == pins exact versions, >= allows upgrades within a range. Always commit to version control.
# Generate
pip freeze > requirements.txt
# Install
pip install -r requirements.txt
# Pin versions
flask==2.3.3
requests>=2.28.0,<3.0.0Poetry
Poetry is a modern dependency manager. pyproject.toml replaces requirements.txt. Virtual environments are managed automatically.
# Initialize
poetry init
# Add dependency
poetry add flask
poetry add pytest --group dev
# Install all
poetry install
# Run command
poetry run python app.pypipenv
pipenv combines pip and virtualenv. Pipfile declares dependencies, Pipfile.lock pin exact versions. --dev separates development dependencies.
# Create environment
pipenv install
# Add package
pipenv install requests
pipenv install pytest --dev
# Activate shell
pipenv shell
# Run command
pipenv run python app.pyConda Environments
Conda manages both Python and non-Python dependencies. environment.yml captures the full environment. Ideal for data science with binary dependencies.
# Create
conda create -n myenv python=3.11
# Activate
conda activate myenv
# Export
conda env export > environment.yml
# Recreate
conda env create -f environment.ymlpip Advanced Usage
Install from Git
Install packages directly from Git repositories. Useful for unreleased versions, forks, or private packages. @branch or @commit pins to a specific version.
# From GitHub
pip install git+https://github.com/user/repo.git
# Specific branch
pip install git+https://github.com/user/repo.git@branch-name
# Specific commit
pip install git+https://github.com/user/repo.git@abc123Editable Install
Editable install (-e) links the package instead of copying. Changes are immediately available without reinstalling. Essential for package development.
# Install in development mode
pip install -e .
# From a specific path
pip install -e /path/to/package
# With extras
pip install -e ".[dev,test]"Constraints & Hashes
Constraints limit which versions can be installed. Hash checking verifies package integrity, preventing supply chain attacks.
# constraints.txt
flask==2.3.3
pip install -c constraints.txt flask
# Hash checking
pip install --require-hashes -r requirements.txtCache Management
pip caches downloaded wheels. --no-cache-dir forces fresh downloads. Purge frees disk space when the cache grows too large.
# Show cache info
pip cache info
# List cached packages
pip cache list
# Purge entire cache
pip cache purge
# Install with no cache
pip install --no-cache-dir flaskCustom Index
--index-url specifies a custom package repository. --extra-index-url adds a fallback. --trusted-host bypasses SSL for internal registries.
# Use custom index
pip install --index-url https://pypi.custom.com/simple/ flask
# Extra index (fallback)
pip install --extra-index-url https://pypi.custom.com/simple/ flask
# Trusted host (no SSL)
pip install --trusted-host pypi.custom.com flaskType Checking (mypy)
Basic Type Hints
Type hints annotate function parameters and return types. Python 3.9+ allows built-in types directly. Hints enable static analysis with mypy.
def greet(name: str, times: int = 1) -> str:
return (f"Hello, {name}! " * times).strip()
def process(data: list[int]) -> dict[str, int]:
return {str(x): x for x in data}Optional and Union
Optional[X] is equivalent to X | None (Python 3.10+). Union types allow multiple possible types. mypy checks all code paths handle all types.
from typing import Optional
def find(items: list[int], target: int) -> int | None:
for i, v in enumerate(items):
if v == target: return i
return NoneGeneric Types
Generics create reusable type-safe containers. TypeVar defines a type variable, Generic makes the class generic. mypy ensures type consistency.
from typing import TypeVar, Generic
T = TypeVar('T')
class Stack(Generic[T]):
def __init__(self) -> None:
self._items: list[T] = []
def push(self, item: T) -> None:
self._items.append(item)
def pop(self) -> T:
return self._items.pop()Protocol
Protocol defines structural subtyping (duck typing with type checking). Any class with the required methods satisfies the protocol, no inheritance needed.
from typing import Protocol
class Closeable(Protocol):
def close(self) -> None: ...
def cleanup(resource: Closeable) -> None:
resource.close()
class File:
def close(self) -> None: print("Closed")
cleanup(File()) # OK - File has close()mypy Configuration
mypy.ini configures type checking strictness. strict enables all checks. Per-module overrides relax rules for tests or legacy code.
# mypy.ini
[mypy]
python_version = 3.11
strict = True
warn_return_any = True
disallow_untyped_defs = True
[mypy-tests.*]
ignore_errors = TruePerformance Tips
List vs Generator
Lists store all elements in memory; generators produce values on demand. Use generators for large sequences iterated once.
# List: all in memory
squares = [x**2 for x in range(1000000)]
# Generator: lazy evaluation
squares_gen = (x**2 for x in range(1000000))
import sys
print(sys.getsizeof(squares)) # ~8MB
print(sys.getsizeof(squares_gen)) # ~200 bytesString Concatenation
String concatenation with += is O(n^2). join is O(n). f-strings are the fastest interpolation method.
# Slow: creates intermediate strings
result = ""
for s in parts:
result += s
# Fast: join in one operation
result = "".join(parts)
# Fast: f-strings
msg = f"Hello, {name}!"Local Variables
Local variable lookups are faster than global or attribute lookups. Assigning frequently-used functions to locals speeds up loops.
import math
# Slow: global lookup
def compute_slow(values):
return [math.sqrt(v) for v in values]
# Fast: local reference
def compute_fast(values):
sqrt = math.sqrt
return [sqrt(v) for v in values]__slots__
__slots__ prevents __dict__ creation, saving 40-50% memory per instance. Significant when creating millions of objects. Cannot add unlisted attributes.
class Point:
__slots__ = ('x', 'y')
def __init__(self, x, y):
self.x = x
self.y = y
p = Point(1, 2)
# p.z = 3 # AttributeErrortimeit & cProfile
timeit measures execution time of small snippets. cProfile shows where time is spent. Use profiling before optimizing to find actual bottlenecks.
import timeit
t = timeit.timeit('sum(range(100))', number=10000)
print(f"{t:.4f}s")
import cProfile
cProfile.run('sum(x**2 for x in range(10000))')Common Pitfalls
Mutable Default Args
Default argument values are evaluated once at definition time. Mutable defaults are shared across all calls. Always use None as the default.
# BUG: default list is shared
def add_item(item, lst=[]):
lst.append(item)
return lst
print(add_item(1)) # [1]
print(add_item(2)) # [1, 2]!
# FIX: use None
def add_item(item, lst=None):
if lst is None: lst = []
lst.append(item)
return lstLate Binding Closures
Closures capture variables by reference. By the time lambdas are called, the loop variable has its final value. Default arguments capture current value.
# BUG: all print 2
funcs = [lambda: i for i in range(3)]
print([f() for f in funcs]) # [2, 2, 2]
# FIX: default argument
funcs = [lambda i=i: i for i in range(3)]
print([f() for f in funcs]) # [0, 1, 2]Integer Caching
Python caches small integers. is checks identity, == checks equality. Never use is for value comparison; use is only for None, True, False.
# Small integers cached (-5 to 256)
a = 256; b = 256
print(a is b) # True (cached)
c = 257; d = 257
print(c is d) # False (not cached)
print(c == d) # Trueis vs ==
is checks whether two references point to the same object. == checks whether two objects have the same value. Use is only for None, True, False.
a = [1, 2, 3]; b = [1, 2, 3]
print(a == b) # True (same values)
print(a is b) # False (different objects)
# Correct usage of is
if x is None: ...
if x is not None: ...GIL
The GIL allows only one thread to execute Python bytecode at a time. Threading is effective for I/O-bound tasks. Use multiprocessing for CPU-bound parallelism.
# GIL prevents true parallelism for CPU-bound tasks
import threading
def cpu_work():
total = sum(i**2 for i in range(10**7))
# Use multiprocessing for CPU work
from multiprocessing import Pool
with Pool(4) as p:
p.map(cpu_work, range(4))Related Python snippets
Copy-paste ready code for common tasks.
Sort Dictionary by Value
Sort a Python dictionary by its values in descending order.
List Comprehension
Quickly generate lists using list comprehensions.
Dictionary Merging
Multiple ways to merge dictionaries.
File Read/Write
Various ways to read and write files.
CSV Processing
Read and write CSV files using the csv module.
JSON Processing
JSON serialization and deserialization.
Regex Matching
Perform regex matching using the re module.
Date Handling
Handle dates and times with datetime.
Decorators
Define and use decorators.
Generators
Save memory using generators.
Context Manager
Custom context managers.
Exception Handling
Complete exception handling mechanism.
Class Inheritance
Class inheritance and method overriding.
Multithreading
Implement multithreading using the threading module.
Multiprocessing
Achieve true parallelism with multiprocessing.
asyncio Asynchronous Programming
Implement asynchronous concurrency with asyncio.
Socket Programming
TCP Socket server and client.
HTTP Requests
Send HTTP requests using the requests library.
Database Operations
Operate on databases using sqlite3.
Virtual Environment
Create and manage Python virtual environments.
pip Install
Common pip package management commands.
Environment Variables
Read and set environment variables.
Logging
Configure and use the logging module.
Unit Testing
Write unit tests using unittest.
Type Hints
Improve code readability with type annotations.
Dataclass
Simplify class definitions with dataclass.
Enum
Define enum types using Enum.
Property Decorator
Control attribute access with property.
Magic Methods
Common magic method examples.
Iterator
Custom iterator implementation.
Coroutine
Basic usage of coroutines.
Was this helpful?