01 · One function, one meaningful job
Split business decisions from external side effects.
validate + calculate + save + print
Remember: split different responsibilities, not every line.
Read function design →Less memorizing. More recognizing. A visual, searchable reference to the essential Clean Code principles you use every day.
These habits reinforce multiple rules from your notes.
Seven related questions. Use these as mental triggers, not as another set of rules to memorize.
Search for a principle or filter by theme. Every row connects a good habit to a code smell.
| Topic | Do this | Watch out for |
|---|
Eight diagrams for concepts where relationships, choices, and execution order matter more than long paragraphs. Each links to its original detailed notes.
Split business decisions from external side effects.
Remember: split different responsibilities, not every line.
Read function design →Reject invalid cases early. Keep the successful path visible.
Remember: guard clauses reduce nesting without changing intended behavior.
Read control flow →Untrusted data gets checked before important side effects.
Remember: catch only what this layer can handle meaningfully.
Read validation and errors →A thin coordinator connects business logic and external systems.
Remember: keep business rules independent of technology. Layers are conceptual, not mandatory folders.
Read separation of concerns →Ask what the data represents before inventing a class.
Remember: stronger modeling should prevent a real mistake, not create ceremony.
Read typing and models →Focus most tests on fast, stable behavior checks.
Remember: verify observable outcomes rather than implementation details.
Read testing →Extract repeated knowledge when rules truly mean the same thing.
Remember: remove duplicate knowledge, not every matching line.
Read DRY and abstraction →Improve structure without intentionally changing observed behavior.
Remember: adding validation changes behavior; keep that separate from structural refactors.
Read refactoring →Open a topic to see the practical rules and one question to ask yourself. Matches your notes' progression. Open the detailed library below for the full explanations and examples.
Keep the quick overview short. Expand a topic, then only the specific rule you need. All substantive text, examples, cautions, and final checklists from your supplied notes are retained below. Original code blocks are reproduced as written; indentation may be missing from the text attachment.
A well-designed function should be:
A function should perform one meaningful job.
def calculate_subtotal(unit_price, quantity):
return unit_price * quantityAvoid combining input, calculations, file writing, database operations, and printing in one function.
A helpful question is:
Can I describe the function without using the word “and”?
If not, it may be doing too much.
Function names should normally start with an action:
calculate_total()
find_customer()
validate_email_address()
save_order()
format_report()Boolean functions should sound like questions:
is_valid_quantity()
has_permission()
can_access_account()Avoid vague names such as:
process()
handle()
check()
do_stuff()There is no strict line limit. Split a function when it:
Do not create separate functions for every individual line. Each function should represent a meaningful operation.
High-level functions should read like workflows:
def register_customer(customer):
validate_customer(customer)
save_customer(customer)
send_welcome_email(customer)Low-level details, such as database statements, should be moved into lower-level functions.
Prefer a small number of meaningful parameters:
def calculate_total(subtotal, tax_rate):
return subtotal * (1 + tax_rate)Use keyword arguments when calls might otherwise be unclear:
calculate_total(
subtotal=100.0,
tax_rate=0.21,
)Group parameters into a data class only when they represent a genuine concept, such as Customer or Product.
Boolean flags can hide multiple behaviors:
generate_report(data, save_to_file=True)Separate functions may be clearer:
report = generate_report(data)
save_report(report, file_path)A Boolean containing actual business data, such as is_premium_customer, is normally fine.
Avoid returning unrelated types:
# Avoid: returns a string or number
def calculate_total(items):
if not items:
return "No items"return sum(items)
Choose a consistent contract:
def calculate_total(items):
if not items:
return 0.0return sum(items)
Alternatively, raise an exception if empty input is invalid.
A query returns information:
def calculate_balance(transactions):
return sum(transaction.amount for transaction in transactions)A command changes state:
def save_customer(customer):
customer_repository.save(customer)Avoid hiding state changes inside functions named get_, find_, or calculate_.
Avoid hidden dependencies on changing global values:
def calculate_total(subtotal, tax_rate):
return subtotal * (1 + tax_rate)Passing the dependency explicitly makes the function easier to understand and test.
A module-level constant is fine when the value is genuinely fixed.
Common side effects include:
If a function modifies its input, make that behavior clear:
def apply_discount_in_place(items, discount_rate):
...Prefer returning a new result when mutation is unnecessary.
Core logic should usually return a result:
def calculate_total(price, quantity):
return price * quantityThe caller can then decide whether to print, save, or send it:
total = calculate_total(20.0, 3)
print(total)This keeps business logic reusable and testable.
Handle invalid and special cases early:
def calculate_discounted_total(
subtotal,
is_premium_customer,
):
if subtotal <= 0:
raise ValueError("Subtotal must be positive.")if not is_premium_customer:
return subtotal
return subtotal * 0.9
Guard clauses reduce nesting and keep the main execution path clear.
Quick Checklist
Before completing a function, ask:
[ ] Does the name clearly describe the purpose?
[ ] Does the function have one responsibility?
[ ] Does it operate at one level of abstraction?
[ ] Are its parameters necessary and understandable?
[ ] Are dependencies explicit?
[ ] Are side effects and mutations obvious?
[ ] Does it return a predictable type?
[ ] Does it avoid unnecessary printing and file access?
[ ] Can it be tested independently?
[ ] Will another developer understand it quickly?
Main takeawayA clean function has a clear name, one meaningful responsibility, explicit inputs, predictable outputs, and no surprising side effects.
The next clean-code topic is simple control flow, covering guard clauses, early returns, nesting, complex conditions, loops, and readable comprehensions.
Handle invalid and special cases early to reduce nesting.
def calculate_discount(user, subtotal):
if user is None:
return 0.0if not user.is_active:
return 0.0
if subtotal <= 0:
return 0.0
return subtotal * 0.10
``
Guard clauses are useful for missing data, invalid arguments, unsupported states, and permission checks.
After return, raise, break, or continue, an else is usually unnecessary.
def get_status(is_active):
if is_active:
return "active"return "inactive"
This keeps the code flatter and easier to scan.
Prefer clear, positive conditions:
if user.is_active:
grant_access()Avoid double negatives:
if not user.is_inactive:
grant_access()Negative conditions are still useful in guard clauses:
if not user.has_permission:
raise PermissionError("Access denied.")Move complicated business rules into descriptive Boolean functions:
def can_place_order(customer):
return (
customer.is_active
and customer.age >= 18
and customer.email_is_verified
and not customer.is_suspended
)Usage becomes easy to understand:
if can_place_order(customer):
approve_order()if user.is_active:
process_user(user)Avoid:
if user.is_active == True:
process_user(user)Use explicit checks only if True, False, and None have different meanings.
For empty collections:
if not orders:
return []If None and an empty collection have different meanings, handle them separately:
if orders is None:
raise ValueError("Orders were not loaded.")if not orders:
return []
Use chained comparisons:
if 18 <= age <= 65:
approve_application()Use membership checks instead of repeated comparisons:
if status in {"pending", "processing"}:
monitor_order()Avoid this common bug:
if status == "pending" or "processing":
...The condition is always true because "processing" is a non-empty string.
Avoid unnecessary indexing:
for customer in customers:
print(customer.name)Use enumerate() when you need the position:
for position, customer in enumerate(customers, start=1):
print(f"{position}. {customer.name}")Use zip() for related collections:
for product_name, product_price in zip(
product_names,
product_prices,
strict=True,
):
print(product_name, product_price)Use continue to skip invalid loop items and keep the main action visible:
for order in orders:
if not order.is_active:
continueif order.quantity <= 0:
continue
process_order(order)
When searching for one item, return as soon as it is found:
def find_customer(customers, customer_id):
for customer in customers:
if customer.id == customer_id:
return customerreturn None
Use any() when at least one item must match:
has_active_customer = any(
customer.is_active
for customer in customers
)Use all() when every item must match:
all_orders_paid = all(
order.is_paid
for order in orders
)Remember:
any([]) # False
all([]) # True
`Check whether that empty-collection behavior matches the business requirement.
Use comprehensions for clear filtering and transformation:
active_customer_names = [
customer.name
for customer in customers
if customer.is_active
]Use a normal loop when the logic contains:
Do not use comprehensions only for side effects:
# Avoid
[send_email(customer) for customer in customers]Use:
for customer in customers:
send_email(customer)A simple conditional expression is fine:
status = "active" if user.is_active else "inactive"For multiple conditions, use a normal function:
def get_user_status(user):
if user.is_premium:
return "premium"if user.is_active:
return "active"
return "inactive"
Instead of repeated branches:
STATUS_MESSAGES = {
"pending": "The order is waiting.",
"processing": "The order is being processed.",
"completed": "The order is complete.",
}def get_status_message(status):
return STATUS_MESSAGES.get(status, "Unknown status.")
Use mappings when one known value directly selects another value. Keep conditionals when each branch performs different actions.
Quick Checklist
[ ] Are invalid cases handled early?
[ ] Is nesting kept shallow?
[ ] Can unnecessary else blocks be removed?
[ ] Are conditions direct and understandable?
[ ] Do complex conditions have meaningful names?
[ ] Are Booleans checked directly?
[ ] Is truthiness used correctly?
[ ] Are range and membership checks concise?
[ ] Do loops iterate directly over objects?
[ ] Would enumerate(), zip(), any(), or all() help?
[ ] Are comprehensions short and free from side effects?
[ ] Are ternary expressions easy to understand?
[ ] Would a mapping be clearer than repeated branches?
Main takeawayClean control flow handles invalid cases early, minimizes nesting, and keeps the normal execution path easy to see.
Validation checks whether data follows your application’s rules:
if quantity <= 0:
raise ValueError("Quantity must be greater than zero.")Error handling decides how the application responds when an operation fails:
try:
quantity = int(user_input)
except ValueError:
print("Enter a valid whole number.")Validate untrusted data when it enters your application, including:
def parse_quantity(raw_quantity: str) -> int:
try:
quantity = int(raw_quantity)
except ValueError as error:
raise ValueError(
"Quantity must be a whole number."
) from errorreturn quantity
Core business functions should also protect critical rules that must never be bypassed.
Validate inputs before calculations, file writes, database updates, or other side effects:
def calculate_total(
unit_price: float,
quantity: int,
) -> float:
if unit_price < 0:
raise ValueError("Unit price cannot be negative.")return unit_price * quantity
Invalid data should not be allowed to travel deeper into the application.
Avoid vague messages:
raise ValueError("Invalid input.")Prefer messages that identify the problem and expectation:
raise ValueError(
"Discount rate must be between 0 and 1."
)A useful message explains:
Do not expose credentials or sensitive internal details.
Avoid broad or bare exception handlers:
try:
quantity = int(user_input)
except:
print("Something went wrong.")Catch only the expected exception:
try:
quantity = int(user_input)
except ValueError:
print("Quantity must be a whole number.")This prevents unrelated programming errors from being hidden.
Only include the operation expected to fail:
try:
quantity = int(user_input)
except ValueError:
print("Quantity must be a whole number.")
else:
total = calculate_total(price, quantity)
save_order(total)A small try block makes it clear which operation the handler belongs to.
A function should catch an exception only when it can:
If it cannot do one of these, let the exception propagate.
Never silently ignore important failures:
# Avoid
try:
save_customer(customer)
except DatabaseError:
passWhen converting a technical exception into a clearer application exception, preserve the original cause:
class InvalidQuantityError(ValueError):
"""Raised when an order quantity is invalid."""The application receives a meaningful error while developers retain the original debugging information.
Built-in exceptions are usually enough for small programs:
Use custom exceptions when callers need to distinguish domain-specific failures:
class OrderError(Exception):
"""Base exception for order failures."""class InvalidOrderError(OrderError):
"""Raised when an order violates a business rule."""
class OrderNotFoundError(OrderError):
"""Raised when an order cannot be found."""
Do not create a separate custom exception for every minor failure.
Return None when absence is normal and expected:
def find_customer(customers, customer_id):
for customer in customers:
if customer.id == customer_id:
return customerreturn None
Raise an exception when the requested operation cannot proceed:
def get_customer(customers, customer_id):
customer = find_customer(customers, customer_id)return customer
A useful convention is:
find_... may return None.
get_... expects the result to exist and may raise an exception.
If a missing dictionary value is expected, use .get():
customer = customers.get(customer_id)If no orders is a normal condition:
if not orders:
return 0.0Exceptions should represent situations that prevent normal completion, not every alternative outcome.
Use context managers for files and similar resources:
with open(file_path, encoding="utf-8") as order_file:
order = json.load(order_file)Use finally when cleanup must always happen and no context manager is available:
connection = create_connection()A validation function should normally either return a valid value or raise an exception:
def validate_discount_rate(
discount_rate: float,
) -> float:
if not 0 <= discount_rate <= 1:
raise ValueError(
"Discount rate must be between 0 and 1."
)return discount_rate
Use naming to clarify behavior:
validate_order() raises when invalid.
is_valid_order() returns True or False.
Users need a clear message:
The order could not be saved. Please try again.Developers need technical details in logs:
try:
save_order(order)
except DatabaseError:
logger.exception("Failed to save order %s", order.id)
display_error(
"The order could not be saved. Please try again."
)Avoid showing users tracebacks, database details, credentials, or internal paths.
Quick Checklist
[ ] Is external input validated at the boundary?
[ ] Are important business rules protected?
[ ] Does validation happen before side effects?
[ ] Are error messages specific and actionable?
[ ] Are only expected exceptions caught?
[ ] Are try blocks small and focused?
[ ] Can this layer genuinely handle the error?
[ ] Is exception chaining used when translating errors?
[ ] Is the choice between None and an exception intentional?
[ ] Are resources cleaned up safely?
[ ] Are user messages separated from technical details?
[ ] Are failures prevented from continuing silently?
Main takeawayValidate data before using it, raise precise exceptions when processing cannot continue, and catch errors only where you can recover or report them meaningfully.
Separation of concerns means keeping different responsibilities in different parts of your application.
Common concerns include:
Avoid mixing input(), calculations, file writing, and print() in one function.
def calculate_order_total(
unit_price: float,
quantity: int,
) -> float:
if unit_price < 0:
raise ValueError("Unit price cannot be negative.")if quantity <= 0:
raise ValueError("Quantity must be positive.")
return unit_price * quantity
The application boundary handles communication with the user:
def main() -> None:
unit_price = float(input("Unit price: "))
quantity = int(input("Quantity: "))total = calculate_order_total(unit_price, quantity)
print(f"Order total: €{total:.2f}")
The calculation can now be reused by a console application, website, API, or test.
A small application commonly has four conceptual layers:
Presentation
Receives input and presents output:
def display_total(total: float) -> None:
print(f"Order total: €{total:.2f}")
ApplicationCoordinates a use case:
def create_order(
unit_price,
quantity,
repository,
):
total = calculate_order_total(
unit_price,
quantity,
)
repository.save(total)return total
Domain
Contains calculations and business rules:
def calculate_order_total(
unit_price,
quantity,
):
return unit_price * quantity
InfrastructureCommunicates with external systems:
def save_total(total, file_path):
with open(
file_path,
"a",
encoding="utf-8",
) as output_file:
output_file.write(f"{total:.2f}\n")You do not always need separate folders or classes for every layer. The important part is keeping their responsibilities distinct.
Keep business results in their useful data type:
def calculate_total(
price: float,
quantity: int,
) -> float:
return price * quantityFormat the result separately:
def format_currency(
amount: float,
symbol: str = "€",
) -> str:
return f"{symbol}{amount:.2f}"Do not return formatted text from a calculation if callers may need the numeric value later.
Business rules should not directly depend on:
Instead of combining calculation and storage:
total = calculate_order_total(items)
save_order_total(total, file_path)This makes the calculation testable without creating files or connecting to external systems.
Dependency injection means passing external dependencies into a function instead of secretly constructing them inside it.
def save_order(order, repository) -> None:
repository.save(order)Production can use a database repository:
save_order(order, database_repository)Tests can use an in-memory repository:
save_order(order, in_memory_repository)Inject dependencies when they perform external work, vary by environment, or need safe substitutes during testing.
Do not inject every small pure helper. That creates unnecessary complexity.
A pure function:
def calculate_tax(
subtotal: float,
tax_rate: float,
) -> float:
return subtotal * tax_rateImpure functions interact with the outside world:
def save_invoice(invoice, repository) -> None:
repository.save(invoice)A useful structure is:
This keeps most business logic easy to test.
The application layer should connect the steps without containing all implementation details:
def create_order(
items,
customer,
repository,
):
total = calculate_order_total(
items,
customer.is_premium,
)return total
This function coordinates the use case while calculations and storage remain separate.
Separation of concerns does not mean:
Separate responsibilities when they:
[ ] Is business logic independent of input and output?
[ ] Are calculations separate from formatting?
[ ] Are files, databases, and APIs isolated?
[ ] Does the application layer coordinate rather than contain details?
[ ] Are external dependencies explicit?
[ ] Can business rules be tested without external systems?
[ ] Are side effects kept near the application boundary?
[ ] Can one responsibility change without affecting unrelated code?
[ ] Do abstractions solve a real problem?
Main takeawayKeep business rules at the core, external interactions at the edges, and use a thin application layer to coordinate the workflow.
Automated tests verify that code continues to behave as expected when you change or refactor it. They reduce regressions, document behavior, and improve confidence, but they only check scenarios you actually write.
Functions are easiest to test when they:
def calculate_total(
unit_price: float,
quantity: int,
) -> float:
return unit_price * quantityCore logic should return values. Application-level code can print, save, or send those values.
A clear test usually has three stages:
def test_calculate_total_multiplies_price_by_quantity():
unit_price = 20.0
quantity = 3result = calculate_total(unit_price, quantity)
Tests should verify the function’s public promise, not private variables or the exact internal sequence.
If you refactor the internal code without changing its behavior, good tests should normally continue to pass.
A test name should describe:
def test_calculate_total_returns_zero_for_free_item():
...def test_parse_quantity_rejects_non_numeric_input():
...
def test_premium_customer_receives_discount():
...
Avoid names such as test_1() or test_function().
Cover at least:
def test_calculate_total_multiplies_valid_values():
assert calculate_total(20.0, 3) == 60.0
Python
def test_calculate_total_allows_zero_price():
assert calculate_total(0.0, 3) == 0.0
Python
def test_calculate_total_rejects_zero_quantity():
with pytest.raises(
ValueError,
match="greater than zero",
):
calculate_total(20.0, 0)Each test should normally verify one behavior.
Prefer separate tests for:
Focused tests make failures easier to diagnose.
Tests should not depend on:
Each test should arrange its own data and clean up its own resources.
Use pytest.mark.parametrize when the same behavior should be checked with several inputs:
@pytest.mark.parametrize(
("unit_price", "quantity", "expected"),
[
(10.0, 1, 10.0),
(10.0, 3, 30.0),
(0.0, 5, 0.0),
],
)
def test_calculate_total(
unit_price,
quantity,
expected,
):
assert calculate_total(unit_price, quantity) == expectedUse separate tests when scenarios require different setup or communicate different business rules.
Fixtures provide reusable test data or resources:
@pytest.fixture
def premium_customer():
return Customer(
name="Pranoy",
is_premium=True,
)Fixtures are useful for:
Keep fixtures understandable. Do not hide important test data inside complex fixture chains.
Tests should not require real files, databases, APIs, or email services unless they are deliberate integration tests.
Use simple test doubles:
class InMemoryOrderRepository:
def __init__(self):
self.saved_orders = []def save(self, order):
self.saved_orders.append(order)
Then test the outcome:
def test_create_order_saves_order():
repository = InMemoryOrderRepository()
order = Order(...)create_order(order, repository)
assert repository.saved_orders == [order]
Prefer a small fake implementation when it is clearer than a mock.
Mock external boundaries, not every internal function.
Too many mocks make tests dependent on implementation details and cause them to fail during harmless refactoring.
Prefer checking observable outcomes:
Avoid recreating the implementation inside the test:
# Less useful
expected = price * quantity
assert calculate_total(price, quantity) == expectedPrefer explicit expected values:
assert calculate_total(20.0, 3) == 60.0If the test repeats the same mistake as the production code, it may pass incorrectly.
Coverage shows which lines or branches ran during tests, but a high percentage does not guarantee useful tests.
Use coverage to find untested areas, not as the only quality measure.
pytest --cov=order_appA healthy test suite typically contains:
Unit tests provide fast feedback. Integration and end-to-end tests confirm that components work together.
Quick Testing Checklist
[ ] Does every test verify one behavior?
[ ] Are test names descriptive?
[ ] Are normal, boundary, and invalid cases covered?
[ ] Do tests follow Arrange, Act, Assert?
[ ] Are tests independent?
[ ] Are expected values explicit?
[ ] Are important exceptions tested?
[ ] Are external dependencies isolated?
[ ] Are mocks used only when necessary?
[ ] Do tests focus on public behavior?
[ ] Is coverage used to find gaps rather than chase a percentage?
[ ] Can the test suite run quickly and consistently?
Main takeawayTest public behavior, keep tests focused and independent, isolate external systems, and use the test suite as protection during refactoring.
Type hints communicate what a function accepts and returns:
def calculate_total(
unit_price: float,
quantity: int,
) -> float:
return unit_price * quantityType hints improve readability, editor support, and static analysis. They do not enforce types at runtime.
Prioritize type hints for:
Local variables usually do not need annotations when their type is obvious:
subtotal = unit_price * quantity
`Add a local annotation when it improves clarity:
orders_by_id: dict[int, Order] = {}Avoid broad collection annotations:
def calculate_total(prices: list) -> float:
return sum(prices)Specify the element type:
def calculate_total(prices: list[float]) -> float:
return sum(prices)Common examples:
customer_names: list[str]
coordinates: tuple[float, float]
orders_by_id: dict[int, Order]
supported_statuses: set[str]If a function only iterates through values, accept Iterable rather than requiring a list:
from collections.abc import IterableUse Sequence when indexing or length is required.
The input type should describe the capabilities that the function genuinely needs.
If a function may return None, show it in the annotation:
def find_customer(
customers: list[Customer],
customer_id: int,
) -> Customer | None:
...The caller must handle the missing case:
customer = find_customer(customers, customer_id)if customer is None:
return
send_notification(customer)
After the check, the type checker understands that customer is a Customer.
This contract provides little protection:
def process(data: Any) -> Any:
...Prefer a precise contract:
def calculate_order_total(
items: list[OrderItem],
) -> float:
...Use Any mainly for genuinely dynamic data at external boundaries, then convert that data into a structured type as early as possible.
A union is useful when multiple types are genuinely supported:
def normalize_customer_id(
customer_id: int | str,
) -> str:
return str(customer_id)Very broad unions can indicate unclear responsibilities:
str | int | float | list | dict | NoneNormalize external data early so the core application works with predictable types.
Instead of passing loosely defined dictionaries:
def calculate_item_total(item: dict) -> float:
return item["price"] * item["quantity"]Use a data class:
from dataclasses import dataclassThen:
def calculate_item_total(item: OrderItem) -> float:
return item.unit_price * item.quantityUse TypedDict when the data must remain dictionary-shaped, such as near JSON or API boundaries.
Use Literal for a small set of accepted values:
from typing import LiteralUse an Enum when those values are important domain concepts:
from enum import EnumRuntime validation is still required for data received from users, files, or APIs.
A protocol describes required behavior:
from typing import ProtocolA function can then depend on the capability rather than a specific technology:
def create_order(
order: Order,
repository: OrderRepository,
) -> None:
repository.save(order)Protocols are helpful for replaceable boundaries such as repositories, API clients, and notification services. Do not create them for every small helper.
Use -> None when a function performs an action without returning a meaningful result:
def save_order(order: Order) -> None:
repository.save(order)Use Never only when a function cannot complete normally:
from typing import Neverdef fail(message: str) -> Never:
raise RuntimeError(message)
Do not claim a function always returns an object when it can return None.
Incorrect:
def find_customer(customer_id: int) -> Customer:
return customers.get(customer_id)Correct:
def find_customer(
customer_id: int,
) -> Customer | None:
return customers.get(customer_id)Alternatively, raise an exception and preserve the stronger return contract.
A type hint can declare that discount_rate is a float:
def apply_discount(discount_rate: float) -> None:
...It cannot guarantee that the value is valid.
Runtime validation is still required:
if not 0 <= discount_rate <= 1:
raise ValueError(
"Discount rate must be between 0 and 1."
)Remember:
Type hints describe kinds of values.
Validation enforces allowed values and business rules.
Tests verify runtime behavior.
Run ty from the project directory:
ty checkA useful development workflow is:
ruff format .
ruff check .
ty check
pytestEach tool answers a different question:
[ ] Are public function inputs and outputs annotated?
[ ] Are collection element types specified?
[ ] Are optional results marked with | None?
[ ] Is Any avoided where a more precise type is possible?
[ ] Are unions limited to genuinely supported types?
[ ] Do annotations match actual behavior?
[ ] Are fixed choices represented clearly?
[ ] Would a data class clarify structured internal data?
[ ] Would TypedDict clarify dictionary-shaped boundary data?
[ ] Would a protocol clarify an important dependency?
[ ] Are type hints supported by runtime validation?
[ ] Does the type checker pass without unnecessary suppressions?
Main takeawayType hints define what kinds of data move through the program. Validation determines whether the actual values are acceptable.
Clear data modeling means representing business concepts explicitly, so valid data is easy to create and invalid states are difficult to represent.
Use the structure that best represents the data:
Avoid passing loosely structured dictionaries throughout the core application.
Instead of:
item = {
"product_name": "Keyboard",
"unit_price": 75.0,
"quantity": 2,
}Use:
from dataclasses import dataclassThis makes required fields and expected types visible and reduces mistakes caused by misspelled dictionary keys.
Behavior that depends mainly on an object’s data can belong to that object:
@dataclass
class OrderItem:
product_name: str
unit_price: float
quantity: intdef calculate_total(self) -> float:
return self.unit_price * self.quantity
Do not add unrelated responsibilities, such as database access or sending emails, to a simple data model.
Validate important rules when the object is created:
@dataclass
class OrderItem:
product_name: str
unit_price: float
quantity: intif self.unit_price < 0:
raise ValueError("Unit price cannot be negative.")
if self.quantity <= 0:
raise ValueError("Quantity must be positive.")
If an object should never exist in an invalid state, reject invalid data during construction.
Models can normalize consistently formatted values:
@dataclass
class Customer:
name: str
email_address: strOnly normalize when the transformation is supported by clear business rules. Do not silently alter meaningful data.
Use frozen=True for values that should not change after creation:
@dataclass(frozen=True)
class CustomerId:
value: intImmutability is useful for:
Remember that frozen=True is shallow. A frozen model can still contain a mutable list. Use tuples for stronger immutability.
Avoid models that allow contradictory states:
class Order:
is_pending: bool
is_processing: bool
is_completed: boolUse one status:
from enum import EnumAn order can now have exactly one status.
Do not allow unrestricted status changes. Use descriptive methods:
def complete(self) -> None:
if self.status is not OrderStatus.PROCESSING:
raise ValueError(
"Only processing orders can be completed."
)self.status = OrderStatus.COMPLETED
Methods such as start_processing(), complete(), and cancel() make the lifecycle visible and protect the object’s rules.
A plain string or integer may not communicate enough meaning:
@dataclass(frozen=True)
class CustomerId:
value: intThese types prevent accidentally passing an order ID where a customer ID is required.
Value objects are useful when a value:
Do not wrap every primitive automatically.
For exact decimal financial calculations, use Decimal rather than float:
from decimal import Decimalprice = Decimal("19.99")
tax_rate = Decimal("0.21")
Create Decimal values from strings, not floats.
A money model can also store currency and prevent adding different currencies:
@dataclass(frozen=True)
class Money:
amount: Decimal
currency: strWhether negative amounts are valid depends on the domain, since refunds and debts may require them.
Never define a shared mutable default:
# Avoid
items: list[OrderItem] = []Use default_factory:
from dataclasses import fielditems: list[OrderItem] = field(default_factory=list)
Each object then receives its own list.
If the model must control how a collection changes, expose a read-only representation and provide meaningful methods:
@property
def items(self) -> tuple[OrderItem, ...]:
return tuple(self._items)def add_item(self, item: OrderItem) -> None:
self._items.append(item)
This lets the model enforce rules around adding, removing, or changing items.
Do not pass raw JSON or API dictionaries throughout the application.
Convert them into domain models early:
def parse_order(raw_order: dict[str, str]) -> Order:
return Order(
order_id=int(raw_order["id"]),
status=OrderStatus(raw_order["status"]),
)The core application then works with known types and validated values.
Build larger models from smaller concepts:
@dataclass
class Customer:
name: str
email_address: EmailAddress
delivery_address: AddressComposition is often more flexible than creating deep inheritance trees such as Customer, PremiumCustomer, and CorporatePremiumCustomer.
Use inheritance only for a genuine and stable “is-a” relationship.
Not every dictionary needs a class, and not every string needs a value object.
Introduce a model when it:
A model should reduce complexity, not add unnecessary ceremony.
Duplication becomes dangerous when the same business rule or knowledge exists in multiple places.
DRY means Don’t Repeat Yourself. It is intended to prevent multiple sources of truth.
Repeated business rule
def calculate_web_discount(subtotal: float) -> float:
if subtotal >= 100:
return subtotal * 0.10return 0.0
return 0.0
Centralized rule
DISCOUNT_THRESHOLD = 100.0
DISCOUNT_RATE = 0.10return subtotal * DISCOUNT_RATE
Now the threshold and rate have one source of truth.
Two functions can look similar while representing different business concepts:
def validate_customer_name(name: str) -> None:
if not name.strip():
raise ValueError("Customer name cannot be empty.")These rules may evolve differently later. Combining them too early could create unnecessary coupling.
A practical guideline is:
Write the code the first time.
Accept some duplication the second time.
Consider abstraction when the pattern appears a third time.
You can abstract earlier when the duplicated code represents an important business rule or inconsistency would be dangerous.
A good abstraction:
Be cautious when an abstraction has:
Remove repeated knowledge, not every repeated line. A little duplication is often better than the wrong abstraction.
Modules and packages give related responsibilities predictable locations.
A module is normally one Python file.
A package is a directory containing related modules.
order_app/
├── __init__.py
├── models.py
├── pricing.py
├── validation.py
└── main.pyA module should have one clear purpose.
For example, pricing.py might contain:
def calculate_subtotal(items):
...def calculate_tax(subtotal, tax_rate):
...
def calculate_discount(subtotal, discount_rate):
...
def calculate_total(items, tax_rate, discount_rate):
...
All these functions concern pricing.
Avoid placing unrelated functions in vague modules such as:
utils.py
helpers.py
common.py
misc.pyThese names often become dumping grounds.
For a larger application, feature-based organization can be more maintainable:
shop/
├── customers/
│ ├── models.py
│ ├── service.py
│ └── repository.py
├── orders/
│ ├── models.py
│ ├── pricing.py
│ └── service.py
└── payments/
├── client.py
└── service.pyCode that changes together stays together.
main() should assemble dependencies and coordinate the application:
def main() -> int:
settings = load_settings()
repository = create_repository(settings)run_application(repository)
return 0
if __name__ == "__main__":
raise SystemExit(main())
It should not contain all business rules, database queries, and formatting logic.
Core business logic should not directly depend on:
Instead, application code should coordinate business logic and infrastructure.
Circular imports often indicate:
Possible solutions include:
Importing a module should not unexpectedly connect to databases or run workflows.
Avoid:
database = connect_to_database()
customers = load_customers()Prefer:
def create_database_connection():
...Construct the resource explicitly in main().
Main takeaway
Group related behavior, make dependencies visible, keep entry points small, and do not let core business logic depend on external technologies.
Refactoring means improving the internal structure of code without intentionally changing its observable behavior.
Understand the existing behavior.
Add or verify tests.
Make one small structural change.
Run the tests.
Review the improvement.
Repeat.
Useful verification commands are:
ruff format .
ruff check .
ty check
pytestRename unclear identifiers
# Before
def calc(p, q):
return p * q
Python
# After
def calculate_subtotal(
unit_price: float,
quantity: int,
) -> float:
return unit_price * quantityExtract a meaningful function
def calculate_subtotal(items) -> float:
return sum(
item.unit_price * item.quantity
for item in items
)
``Extract code when it represents a business rule, separate responsibility, or reusable concept.
Inline an unnecessary function
If a function only forwards a value and contributes no meaning, remove it.
# Unnecessary wrapper
def get_total(order):
return order.calculate_total()Call the meaningful operation directly:
total = order.calculate_total()
Introduce explanatory variables
Python
subtotal = unit_price * quantity
discount_amount = subtotal * discount_rate
final_total = subtotal - discount_amountThis is often clearer than one compressed expression.
Replace magic values
PREMIUM_DISCOUNT_RATE = 0.10
FREE_DELIVERY_THRESHOLD = 100.0Names should explain the business meaning.
Simplify conditions
Use:
If an operation primarily depends on one object’s state, it may belong on that object:
@dataclass
class OrderItem:
unit_price: float
quantity: intDelete:
Version control preserves history.
Adding validation changes behavior:
if quantity <= 0:
raise ValueError("Quantity must be positive.")That may be valuable, but it is not purely structural refactoring. Keeping these changes separate makes reviews and debugging easier.
Main takeaway
Refactor through small, test-protected changes. Improve structure without redesigning more than the current problem requires.
Object-oriented design combines related state and behavior in cohesive objects.
A class is useful when:
Example:
class BankAccount:
def __init__(self, balance: float = 0.0) -> None:
if balance < 0:
raise ValueError(
"Initial balance cannot be negative."
)self._balance = balance
self._balance += amount
The account controls its own valid state.
Do not create a class for a simple stateless transformation:
def calculate_tax(
subtotal: float,
tax_rate: float,
) -> float:
return subtotal * tax_rateA TaxCalculator class without meaningful state would add unnecessary ceremony.
A class should have one focused responsibility.
Avoid oversized manager classes containing:
These responsibilities should normally be separated.
Encapsulation means controlling how important internal state changes.
Instead of:
order.status = OrderStatus.COMPLETEDPrefer:
def complete(self) -> None:
if self.status is not OrderStatus.PROCESSING:
raise ValueError(
"Only processing orders can be completed."
)self.status = OrderStatus.COMPLETED
Use a property when attribute access requires validation, calculation, or controlled exposure:
@property
def balance(self) -> float:
return self._balanceDo not hide network requests or expensive operations inside properties.
Composition assembles an object from smaller collaborators:
class OrderService:
def __init__(
self,
repository,
payment_service,
notifier,
) -> None:
self._repository = repository
self._payment_service = payment_service
self._notifier = notifierPrefer composition when components vary independently.
Use inheritance only for a genuine and stable “is-a” relationship where subclasses can replace the base type without surprising callers.
Constructors should establish valid state. They should not unexpectedly:
Warning signs include:
Use a class when data, behavior, and lifecycle rules form one cohesive concept. Prefer functions for simple transformations and composition for combining capabilities.
You chose to skip the detailed SOLID lesson, so this section is only a reference.
SOLID is a group of five object-oriented design principles:
S: Single Responsibility Principle
A component should have one focused responsibility or one main reason to change.
O: Open/Closed Principle
Code should be extensible without requiring frequent modification of stable behavior.
L: Liskov Substitution Principle
A subtype should be usable wherever its base type is expected without breaking the promised behavior.
I: Interface Segregation Principle
Prefer small, focused interfaces over large interfaces that force clients to depend on operations they do not use.
D: Dependency Inversion Principle
High-level behavior should depend on meaningful abstractions rather than directly depending on infrastructure details.
SOLID should not be applied mechanically. Overusing it can create:
Use the principles only when they solve actual design problems.
Status
[-] SOLID principles: skippedSecurity starts with treating external data and systems as untrusted.
Validate:
Prefer allowlists:
SUPPORTED_FILE_TYPES = {
".csv",
".json",
".txt",
}Accept only formats the application supports.
Avoid executing external input
Never use eval() on external input:
# Unsafe
result = eval(user_input)Use appropriate parsers such as json.loads().
Protect secrets
Do not place credentials in source code:
import osdef load_api_key() -> str:
api_key = os.getenv("ORDER_API_KEY")
return api_key
Also keep secrets out of:
Never build SQL commands using raw external strings:
cursor.execute(
"SELECT * FROM customers WHERE email = ?",
(email_address,),
)The placeholder syntax depends on the database library.
Set timeouts
External requests must have time limits:
response = client.get(
endpoint,
timeout=10,
)
Retry carefullyRetry only:
Retries can duplicate side effects such as payments unless the operation supports safe duplicate handling.
Use transactions
When multiple updates form one business operation, define how they succeed or fail together.
Examples include:
A dependency failure should not make security decisions more permissive:
def has_permission(user) -> bool:
try:
return load_permissions(user)
except PermissionServiceError:
return False
Separate user errors from technical diagnosticsUsers need safe, understandable messages. Developers need detailed logs.
Do not expose:
Examples:
Test:
Treat external data as untrusted, keep secrets out of code and logs, make failures explicit, and design important operations with their failure paths in mind.
Performance work should be based on evidence, not assumptions.
Recommended sequence
Make the code correct.
Make it clear.
Define a measurable target.
Measure the current performance.
Profile the application.
Optimize the actual bottleneck.
Measure again.
Verify correctness.
Define measurable requirements
Examples:
Process 100,000 orders in under two seconds.
Plain Text
Complete 95% of API requests within 300 milliseconds.
Plain Text
Use less than 500 MB of peak memory.
Measure timeUse timeit for small operations and perf_counter() for application sections:
from time import perf_counterstart_time = perf_counter()
result = process_orders(orders)
elapsed = perf_counter() - start_time
Profile before optimizing
Use cProfile:
python -m cProfile -s cumulative app.pyLook for:
Choose containers based on operations:
For repeated lookup by identifier:
orders_by_id = {
order.order_id: order
for order in orders
}This is often better than repeatedly searching a list.
Move invariant work outside loops
settings = load_settings()for order in orders:
process_order(order, settings)
Do not reload unchanged settings for every order.
Reduce external I/O
Database and network calls are often more expensive than Python calculations.
Consider:
Process large files line by line:
with file_path.open(encoding="utf-8") as data_file:
for line in data_file:
process_line(line)Use batching when the external system supports bulk operations.
Caching is appropriate when:
A cache adds complexity around expiration and invalidation.
CPU-bound work may benefit from better algorithms, optimized libraries, or multiprocessing.
I/O-bound work may benefit from batching, connection reuse, or controlled concurrency.
Concurrency is not automatically faster. It introduces failure, cancellation, ordering, and resource-management concerns.
After every optimization:
First make the code correct and clear. Then measure realistic workloads, optimize the real bottleneck, and verify the improvement without sacrificing correctness.
Design and organization
[ ] Is repeated business knowledge centralized?
[ ] Are abstractions meaningful and stable?
[ ] Does each module have a cohesive responsibility?
[ ] Is the entry point small?
[ ] Are dependencies clear and correctly directed?
Refactoring and objects
Plain Text
[ ] Are refactorings small and protected by tests?
[ ] Are unclear names and complex conditions improved?
[ ] Are classes used only when they add real value?
[ ] Do objects protect their valid state?
[ ] Is composition preferred over unnecessary inheritance?
Security and reliability
Plain Text
[ ] Is external input validated?
[ ] Are secrets kept outside source code and logs?
[ ] Are database queries parameterized?
[ ] Do external calls have timeouts?
[ ] Are retry and transaction policies explicit?
[ ] Are failure paths tested?
Performance
Plain Text
[ ] Is there a measurable performance problem?
[ ] Has the bottleneck been profiled?
[ ] Are appropriate algorithms and data structures used?
[ ] Are repeated database and network calls minimized?
[ ] Are streaming, batching, or caching justified?
[ ] Do tests still pass after optimization?
Updated StatusClean code is not about applying every principle or pattern. It is about choosing structures that make the program clear, correct, secure, testable, maintainable, and appropriately efficient.
Use code smell recognition instead of memorizing isolated rules.
def process(order, db, notify=False):
if order:
if order["quantity"] > 0:
total = order["price"] * order["quantity"]
db.save(total)
if notify:
print("Saved")
return total
return "Invalid order"def calculate_total(price: float, quantity: int) -> float:
if price < 0 or quantity <= 0:
raise ValueError("Invalid price or quantity")
return price * quantity
def create_order(order, repository) -> float:
total = calculate_total(order["price"], order["quantity"])
repository.save(total)
return totalFive questions before you mark a task done.
Tick these as you inspect a change. Progress stays in this browser.
Use a real pull request or a function from your project as the learning material.
Close the notes. Say one principle and one warning sign from memory.
Find one example in code: nesting, hidden I/O, duplicated rules, weak tests.
Make one small, test-protected change, or explain why no change is needed.
Suggested review pattern: day 1 (functions + control flow), day 2 (validation + errors), day 4 (architecture + tests), day 7 (typing + models), day 14 (DRY + OOP + refactoring), day 30 (security + performance). Practice recall before reopening the full notes.